Reviewing AI Rules
Reviewing an AI rule means examining what it accomplishes, whether that purpose still matters, and how well the current way of achieving it works. A rule may be an instruction, a required step, a check or a limit on action. Each can deserve a different response when the work changes.
An improvement in AI capability gives you a reason to look again. So can a new tool, a recurring failure, a change in workload or evidence that a required step creates unnecessary work. The useful result is a better working arrangement: one that delivers what people need while preserving the conditions that matter.
This guide proposes a practical way to conduct that review. It is a reasoning aid, not a method validated across enterprises. It applies Effective Enterprise AI’s focus on durable design for value delivery to the instructions and controls surrounding AI-enabled work.
Start with the job the rule performs
Choose a rule whose cost or usefulness is in question. Describe the problem it addresses before deciding what to do with it.
Useful questions include:
- What result does this rule help someone achieve?
- What failure or unacceptable action is it meant to prevent?
- Who depends on it, including people further along the delivery path?
- What evidence led us to adopt it, and do those conditions still apply?
A rule’s history can help answer these questions. An instruction added after an error may preserve a valuable lesson, but its wording may extend beyond the circumstances that made it useful. Recover the relevant context from the people, records or systems that depend on it. If the rationale remains unclear, record that uncertainty and identify who can authorize further investigation or a limited test. Missing history is not permission to remove a protection.
Keep this investigation proportionate. Prioritize rules whose consequences or costs make them worth examining, rather than documenting every sentence in every prompt. Deciding which rule to investigate is separate from deciding which change is appropriate to test.
Separate the purpose from the way it is achieved
A rule has a function, the job it performs, and a mechanism, the way that job is carried out. A fixed drafting sequence, for example, might help an AI system produce a complete recommendation. Completeness is the desired result; the sequence is one way to pursue it.
The following questions offer several ways to examine a rule. They overlap: a required handoff format might support coordination, evidence and recovery at the same time.
| What the rule serves | Question to ask |
|---|---|
| Purpose and acceptance | Does it make clear what a useful, acceptable result must accomplish? |
| Authority and obligations | Does it limit what may be accessed, changed or promised? |
| Coordination and shared state | Does it keep contributions compatible and prevent conflicting actions? |
| Quality and evidence | Does it establish an adequate basis for accepting the work? |
| Method support | Does it prescribe an approach that helps this AI setup perform the task? |
| Recovery and learning | Does it help people detect, correct or learn from a failure? |
These are practical lenses, not an exhaustive classification. If a rule serves several functions, examine each before changing the wording or removing the step.
Consider this illustrative example. A team requires an AI assistant to produce a prescribed outline before drafting a customer proposal. The outline may help expose missing requirements. A newer setup might produce equally complete drafts using clear acceptance criteria and a check against the customer’s requirements. That possibility is worth testing.
The same workflow may also require conflicting prices to be reconciled before the proposal is sent, and approval for commitments beyond delegated authority. Better drafting does not by itself resolve conflicting commitments or change who may make them. The means of coordinating and enforcing those boundaries can evolve, but the relevant dependencies and authority decisions still need attention.
For each function, ask whether the proposed arrangement will still do the job. The comparison below pairs three purposes with questions to ask about a change.
Choose a specific change to investigate
Reviewing a rule can lead to several useful decisions.
| Option | What would justify it? |
|---|---|
| Retain it | The purpose remains relevant and the rule performs its job at a proportionate cost. |
| Revise it | The purpose is sound, but the wording, scope or trigger creates avoidable problems. |
| Move its function elsewhere | Evidence shows that a tool permission, interface, automated check or other arrangement performs the job more effectively. |
| Test a relaxation | There is a credible alternative and a bounded comparison can reveal relevant losses before broader use. |
| Retire it | Its purpose is no longer needed, or another verified arrangement already serves that purpose adequately. |
A rule may need more than one disposition. For example, separate a prescribed drafting sequence from the requirement to obtain approval for commitments beyond delegated limits. You can test the sequence while retaining the approval boundary.
Moving a requirement out of a prompt does not necessarily eliminate the work it performs. A permission setting may now enforce it. A tool may check it. A person may have inherited the responsibility. Include those changes when comparing effort, cost and reliability.
The same applies to apparent improvements or regressions. Record the model, settings, prompts, tools and review process in use. Anthropic’s investigation of Claude Code quality reports, for example, identified interacting product settings and context-handling changes rather than a change to the underlying API model. The surrounding arrangement matters when interpreting results. Anthropic’s quality investigation
Run a bounded comparison
Reviewing a consequential rule does not mean relaxing it in live work. For a first comparison, choose one method change whose consequences you can contain, in work you are authorized to change. If the change belongs to another decision owner, take the findings to that owner before testing a relaxation.
Draft-only work can provide a useful starting point because outputs can be examined before they affect a customer or an external system. Prevent the trial from sending proposals or changing external records, for example by disabling those actions in its tools. Keep relevant access and commitment boundaries in place.
1. Describe the result and the alternative
State the rule’s purpose and propose one specific alternative. In the proposal example, you might compare the prescribed outline with instructions that define required content and leave the drafting sequence open.
Name the recipient who will use this step’s output and the customer whose outcome the broader delivery flow serves. The recipient might be the person reviewing and combining the proposal; the customer needs a clear, accurate basis for deciding whether the offer meets their needs. When an output goes directly to the customer, that customer is also its immediate recipient. Still examine both the usefulness of the output and what happens when the customer acts on it.
2. Decide what would count as improvement or loss
Set the acceptance and stop conditions before examining the results. Acceptance conditions define an adequate result; stop conditions define when to end the trial early. Include what matters for this workflow: completeness, accuracy, retained qualifications, recipient understanding, rework and relevant limits on action.
A useful proposal must do more than sound convincing. It may need to preserve uncertainty, distinguish an estimate from a commitment, and give the recipient enough understanding to explain or revise the recommendation.
Choose checks that can reveal the failure the rule was meant to prevent. Avoid requiring the old sequence merely because it is familiar. Equally, do not let an attractive final output conceal an unacceptable action used to produce it. Anthropic’s evaluation guidance distinguishes the final state of the task from the recorded actions and interactions that produced it. Both can need assessment, and overly rigid grading can reject valid approaches. Anthropic’s guide to agent evaluations
3. Compare under similar conditions
Use matched cases with the existing rule and the proposed alternative. Keep the model, settings, tools and available information the same so that the comparison can inform the decision about the rule. Include difficult cases relevant to its purpose, rather than only routine successes.
Repeat cases where variation could change your conclusion. Where practical, have the recipient assess outputs without knowing which method produced them. If the format or other clues reveal the method, record that limitation and apply the same acceptance criteria to both. Record the work required to create, review and repair each result.
The scale of the comparison should reflect the consequences of the decision. A small exercise may reveal a problem or support a limited next step. It cannot establish protection against rare, consequential failures or prove that the arrangement will remain effective across future models.
4. Follow the effects beyond this step
An assistant may produce a draft faster while the recipient spends longer checking it. A reviewer may save time while the customer still waits for an unrelated decision. Preserve genuine local benefits, but distinguish them from demonstrated improvement in the customer’s outcome.
Ask what happens after the output leaves this step. Does it reduce ambiguity, waiting or rework? Does it help someone make a better decision? Has necessary work been eliminated, or transferred to someone else?
Some consequences will take longer to observe. Record the local result and what remains uncertain about the customer contribution, then identify how and when to revisit that uncertainty. An immediate score should not stand in for evidence that can only emerge through use.
Record what you learned and when to look again
The decision may be to keep the rule, narrow it, replace it or investigate further. Record enough context that the next person can understand why. Retiring a rule should be an informed conclusion about its function and replacement, rather than a reward for producing a shorter instruction file.
If results are inconclusive, record what the comparison could not establish. Keep required protections and decide whether better checks or different cases would make a further test useful. No detected difference does not prove equivalence.
Identify who can authorize the change and who will respond if the revised arrangement stops working. Preserve a recoverable baseline where recovery is possible. Where an action has irreversible consequences, the design needs protection before that action occurs.
A useful review trigger might be a change in model, tool, workload, authority or observed failure. Choose triggers that bear on the rule’s job. The aim is to keep the reasoning available when circumstances change, without making rule maintenance a larger burden than the problem it addresses.
A short record to copy
- Rule and functions: What does it require? Which results, protections or dependencies does each part serve?
- People served: Who receives this output? How should it contribute to the customer’s outcome?
- Ownership and authority: Who can decide on this change, and who will respond to problems?
- Proposed alternative: What will change, and what conditions will remain the same?
- Acceptance and stop conditions: What evidence would support the alternative or reveal a loss?
- Observed result: What happened locally and downstream? What remains uncertain?
- Decision and follow-up: What will be retained or changed, how can problems be contained or corrected, and what will trigger another review?
Evidence behind this approach
This guide proposes a way to reason about rule changes. Its functions and comparison exercise have not been validated as a general method across enterprises.
Research and engineering experience support examining instructions in context. In coding-agent benchmarks, Evaluating AGENTS.md found no significant aggregate task-success improvement from context files over the no-context condition, while developer-provided guidance outperformed generated guidance. Those results do not show that instructions are generally unnecessary or that shorter files are inherently better. Evaluating AGENTS.md, June 2026 revision
The distinction between local gains and wider change also matters. In Shifting Work Patterns with Generative AI, researchers from Microsoft Research and Harvard Business School studied randomized access to Microsoft 365 Copilot among knowledge workers. They found individual time savings without detecting changes in the quantity or composition of tasks. That does not establish an absence of enterprise benefit; it illustrates why a particular measured gain cannot answer every question about work and value. Shifting Work Patterns with Generative AI, November 2025 revision
Related resources
- Purpose, boundaries and the people served: distinguish the immediate recipient from the customer whose outcome the delivery flow serves.
- Evidence and feedback: connect acceptable work with evidence of its contribution to customer outcomes.
- Vision, path and means: place a bounded change within the broader design of effective AI-enabled work.