Revision and change control
| Revision | Date | Change | Author | Reviewer and approver |
|---|---|---|---|---|
| 1 | 3 Aug 2026 | Original SOP | GPT-5.6 Sol | Viraj Ganguli |
| 2 | 5 Aug 2026 | Revision 2 | GPT-5.6 Sol | Tory Clasen |
- Standard Operating Procedure: AI Agent Evaluation
This standard operating procedure (SOP) defines a practical baseline for evaluating artificial intelligence (AI) agents. It helps teams make sound selection, release, and monitoring decisions.
The SOP applies when an AI system plans work, uses tools, changes data, or acts for a user. It covers experiments, comparisons, regression tests, release tests, and production monitoring.
Must marks a required control. Should marks a recommended control. May marks an optional method.
Teams have freedom to choose tools, test methods, metrics, and document formats. Teams may combine records or automate steps when the result meets this SOP.
Every decision evaluation must follow these rules:
- State the decision before building the evaluation.
- Test the complete agent system, not only the model.
- Use cases that represent real work and important failures.
- Set success criteria before the final run.
- Protect government, customer, company, and personal data.
- Keep enough evidence to explain and repeat the result.
- Match the review effort to the risk.
- Keep exploration separate from decision evidence.
Teams may use fast experiments during development. An experiment becomes decision evidence only after the team meets the applicable controls in this SOP.
Each evaluation must name these roles:
- Decision owner: Makes the selection, release, or deployment decision.
- Evaluation owner: Plans the evaluation and maintains its evidence.
- Technical reviewer: Checks cases, graders, results, and conclusions.
- Security reviewer: Checks data, access, tools, and contract controls when required.
One person may hold more than one role. For an Elevated evaluation, the main developer must not be the only reviewer and decision owner.
The team may ask a domain expert to review specialized tasks. The team does not need a standing evaluation committee.
The evaluation owner must select a level before the final run.
Use Routine when all these conditions apply:
- The work is internal or occurs in an isolated test environment.
- The data is synthetic, public, or approved for the test environment.
- A person checks outputs before any material action.
- Errors are easy to detect and reverse.
- The agent has limited tools and permissions.
The decision owner and technical reviewer approve a Routine evaluation.
Use Elevated when any of these conditions apply:
- The agent uses controlled government, customer, personal, or company data.
- The agent can change production systems or send external communications.
- An error can cause legal, security, privacy, financial, mission, or safety harm.
- The agent has broad tools, network access, or limited human oversight.
- A contract, customer, security plan, or company policy requires more review.
The decision owner, technical reviewer, and security reviewer approve an Elevated evaluation. Add legal, privacy, contracts, or customer review when an applicable requirement calls for that review.
When the correct level is unclear, use Elevated until the decision owner records a lower classification.
Reassess the level when data, users, tools, permissions, autonomy, scale, or consequences change.
Use the 7 steps below for decision evaluations. Teams may combine meetings and records.
flowchart TD
A[1. Define the decision] --> B[2. Design representative cases]
B --> C[3. Choose measures and graders]
C --> D[4. Prepare a controlled run]
D --> E[5. Run the evaluation]
E --> F[6. Analyze and decide]
F --> G[Record the decision and retain evidence]
G --> H{Deploy or continue?}
H -->|Yes| I[7. Monitor and improve]
H -->|No| J[Close the evaluation]
I -->|Material change| A
Create a short evaluation plan before the final run. The plan must state:
- the decision and decision owner;
- the agent version and intended use;
- the users, environment, data, tools, and permissions in scope;
- the baseline or minimum acceptable result;
- the important failure types;
- the success criteria;
- the evaluation level;
- the planned cost and schedule; and
- the limits of any conclusion.
Use a measurable claim when comparing systems. For example:
The candidate completes the approved tasks at least as accurately as the current system. It has no critical security failures.
Do not use claims such as the agent is good or the agent passed without defined criteria.
Lesson learned — begin with the claim. A large test suite can still give an unclear result. Define whether the evaluation tests quality, speed, cost, safety, or a stated combination. Convert broad claims into measurable criteria and acceptable trade-offs before building cases.
Map the agent's main tasks, tool actions, outputs, handoffs, and recovery paths. Select cases that support the decision and match the evaluation level.
Verifier's Rule states that AI is easier to train on tasks that are easier to verify. For an evaluation program, the rule is a task-selection guide, not proof that a task is valuable, safe, or ready for automation.
A task is easier to verify when:
- people agree on what a good result is;
- the team can check a result quickly;
- the team can check many results at the same time;
- the check closely tracks the quality that matters; and
- the check gives useful differences between weak and strong results.
The evaluation owner should use these properties to plan the company's evaluation portfolio. Prioritize tasks that combine business value, material risk, and reliable verification. These tasks support faster tests, clearer comparisons, and more frequent regression runs.
Do not select tasks only because they are easy to grade. Easy-to-grade tasks can create a coverage bias when important work requires expert judgment or delayed outcomes. Include material tasks with difficult verification, or record them as coverage gaps. Use expert review, production signals, longer observation periods, or narrower claims for those tasks.
Scope each task so the verifier can judge a meaningful work result. Define the starting state, allowed actions, required end state, accepted alternatives, critical failures, and evidence that the verifier can inspect. Add answer keys, tests, schemas, state checks, or audit records when these controls improve verification.
Split a broad workflow into smaller tasks when each smaller result has independent value and a clear check. Keep end-to-end cases for important handoffs and combined failures. Do not claim that the agent can complete the broad workflow when the evaluation verifies only its parts.
Limit every conclusion to the qualities that the verifier checks reliably. If the verifier is slow, noisy, or incomplete, narrow the task or add human review. If neither action produces reliable evidence, do not use the task for a quantitative approval threshold.
Include these case types when they apply:
- normal work;
- valid edge conditions;
- unsafe or prohibited requests;
- malicious inputs, including prompt injection;
- tool, service, and network failures;
- harmless inputs that should not cause action; and
- cases that require refusal or human help.
Use realistic cases without placing protected data in an unapproved system. Label synthetic data and record the source of other data.
Each case must have:
- a stable identifier;
- a purpose;
- an input or fixture;
- an expected result or scoring rule; and
- any critical failure condition.
The team decides how many cases are enough. The evaluation report must describe important coverage gaps.
Lesson learned — balance positives with controls. A suite of known defects can reward an agent that always reports a problem. Include harmless and no-action cases to measure correct restraint and false positive cost. Use difficulty labels and a coverage matrix to show gaps before summary scores hide them.
Choose measures that support the decision. Useful measures can include:
- task completion and correctness;
- security, privacy, safety, or policy failures;
- false approvals and false refusals;
- reliability across repeated runs;
- time, cost, tool calls, and human correction; and
- recovery from failures.
Set required thresholds before the final run. Do not let a high average score hide a critical failure.
Use the simplest valid grader. A grader can use code, rules, model judgment, human judgment, or a combination. Test important graders with correct, incorrect, partial, and malformed outputs.
When judgment affects the decision, have a domain expert or technical reviewer check a useful sample. Allow more than one correct answer when the task permits valid variation.
Lesson learned — grade behavior, not presentation alone. A correct final answer does not prove that the agent followed required controls. When the task requires specific behavior, grade tool calls, command arguments, destinations, file changes, and other recorded actions.
If results can vary between runs, use enough repeated trials to understand that variation. State the trial count and uncertainty in the report.
Before the final run, record the exact test configuration. Include:
- model and provider;
- agent code and instructions;
- tools and permissions;
- cases and graders;
- relevant runtime limits;
- test environment; and
- source revision or version.
Lesson learned — the manifest is the scope. Folder contents can change and do not prove which items the team selected. Use a versioned manifest to identify each approved case, prompt, fixture, grader, model, provider, agent, and trial assignment.
Complete these checks before a paid or decision run:
- Confirm that the agent cannot read expected answers or grader logic.
- Confirm that the test environment contains only approved data.
- Confirm that tools, network access, and credentials follow the applicable security plan.
- Confirm that time, cost, retry, and action limits work as expected.
- Confirm that logs exclude secrets and follow data handling rules.
- Run a small end-to-end test.
Lesson learned — validate cheaply before scaling. Run offline checks for schemas, manifests, fixtures, graders, and evaluation tools before model use. Then run a small end-to-end test before a large paid evaluation. These checks protect the decision budget and separate setup defects from agent behavior.
An Elevated evaluation must use an approved environment and least-privilege access.
Least privilege means that the agent receives only the access needed for the test.
Lesson learned — configuration is not enforcement. A value in a configuration file does not prove that the limit works. Test retry, output, context, destination, action, and cost limits during preparation. Block a test route when the route cannot enforce a required limit.
Lesson learned — per-case isolation matters. Repository access can expose fixtures or expected answers from other cases. Isolate each case from unrelated files and block path traversal or indirect access. Use negative tests to confirm the isolation before the final run.
Stop preparation when a required control does not work. Fix the control or approve an exception before the final run.
Create a unique run identifier. Keep the approved configuration unchanged during the run.
The run must:
- reset state between independent cases unless state is part of the test;
- use only approved tools, data, models, and network destinations;
- capture outputs, grader results, important tool actions, errors, time, and cost;
- keep partial evidence from failed attempts; and
- stop when a security, data, cost, or safety limit is reached.
Do not change thresholds after viewing final results. Do not rerun only failed cases and merge them into the original result.
If a run fails for technical reasons, record the reason. Start a new run after the team fixes the problem.
Separate these failure sources in the report:
- agent or model failure;
- provider or service failure;
- evaluation tool failure;
- expected security or policy block; and
- bad or ambiguous test data.
Do not count an evaluation tool failure as an agent-quality failure. Do not remove difficult results after viewing them.
Lesson learned — keep platform quality separate. One malformed answer can cause several dependent grader failures. Report the root output failure and keep the dependent failure details. Do not create scores for results that the graders could not compute.
The report must include:
- the decision and tested configurations;
- the evaluation level;
- the cases, measures, thresholds, and trial counts;
- results for each required criterion;
- critical failures and failure sources;
- important limits and coverage gaps;
- cost and operational effects when relevant;
- the evidence location; and
- the recommended decision.
The decision owner records one result:
- approve;
- approve with stated limits;
- hold for more evidence;
- return for changes;
- reject; or
- retire or roll back.
An approval with limits must name each limit, owner, and review date.
For a deployed agent, record:
- the production signals that the team will monitor;
- the person who reviews those signals;
- the conditions that trigger investigation or rollback; and
- the method for reporting user or operator problems.
Add useful production failures to development tests after checking data rights and privacy. Keep decision cases separate when repeated development use could reveal expected answers.
Run relevant regression tests after a material change. Material changes include changes to models, prompts, tools, permissions, graders, data, or workflows.
Review an active Elevated evaluation at least once each year. Also review it after a serious incident or material change.
Lesson learned — ownership completes the benchmark. A test suite remains experimental until named people own its use and maintenance. Assign owners for approval criteria, production feedback, regression cases, result approval, and retirement. Reproducible tests do not replace operational accountability.
Keep one evidence package for each decision evaluation. A folder, ticket, repository, or approved evaluation system can hold the package.
flowchart TD
A[Evaluation plan and level] --> D[Controlled run]
B[Cases and graders] --> D
C[Run configuration] --> D
D --> E[Outputs, actions, errors, time, and cost]
E --> F[Evaluation report and reviews]
F --> G[Decision]
G -->|Deployed agent| H[Monitoring and regression tests]
H -->|New failure or material change| A
H -->|Update cases| B
The package must contain:
- the evaluation plan and level;
- the case and grader versions;
- the run configuration and identifier;
- the raw results or a controlled link to them;
- the evaluation report;
- reviews, approvals, and exceptions; and
- the final decision.
Protect evidence according to its most sensitive content. Follow contract, legal, customer, and company retention rules.
Evidence must support the stated result. Routine evaluations do not need duplicate forms, signatures, or records that do not support the decision.
Lesson learned — retain evidence deliberately. Store generated results in approved evidence storage, separate from committed test fixtures. Record provenance and retention data for large inputs. Use content hashes or approved large-file storage when needed for integrity. Do not use repository history as general storage for test artifacts.
The decision owner may approve a temporary exception when no higher authority is required. An exception record must state:
- the omitted or changed control;
- the reason;
- the resulting risk;
- an alternate control, when available;
- the approving person; and
- the expiration date or closing condition.
An exception cannot override law, contract terms, customer direction, or an approved security plan. Reapprove the exception after a material change.
Stop an affected run when continued operation can expose data, exceed authority, or cause harm. Preserve relevant evidence without increasing the exposure.
Follow the project incident process for security, privacy, safety, or data events. Notify the evaluation owner and decision owner when an event can affect results.
Mark affected results as invalid or limited. Review any decision that used invalid results. For a deployed agent, use the approved containment or rollback process.
Record the cause, correction, and regression test after the immediate response is complete.
Use version control or another approved change record for cases, graders, prompts, and evaluation code.
Require a new decision run when a change can alter the result or its meaning. Examples include changes to:
- the agent, model, prompt, tools, or permissions;
- selected cases or expected results;
- grader logic or thresholds;
- the test environment; or
- the intended users or workflow.
Correcting spelling or explanations does not require a new run when the correction cannot change execution or interpretation. Keep old reports linked to the versions that produced them.
Teams may add fields or combine these templates. Keep every required item somewhere in the evidence package.
# Evaluation Plan: [name]
- Owner:
- Decision owner:
- Technical reviewer:
- Security reviewer, if required:
- Evaluation level and reason:
- Decision:
- Agent version and intended use:
- Users and environment:
- Data, tools, permissions, and limits:
- Baseline or minimum result:
- Important failures:
- Measures and thresholds:
- Cases, graders, and trial count:
- Maximum cost:
- Applicable government or customer requirements:
- Known coverage gaps and limits:
- Approval and date:# Evaluation Report: [name and run identifier]
- Decision tested:
- Configurations and versions:
- Evaluation level:
- Run date:
- Cases and trial count:
- Result for each threshold:
- Critical failures and failure sources:
- Cost and operational findings:
- Coverage gaps and limits:
- Evidence location:
- Reviewer findings:
- Recommended decision:
- Final decision, owner, and date:
- Approval limits, owners, and review dates:Before using an evaluation for a selection, release, or deployment decision, confirm these items:
- The evaluation has a decision owner and an evaluation level.
- The plan defines scope, success criteria, important failures, and conclusion limits.
- Cases represent real work and important risks.
- Graders have checks that match their effect on the decision.
- The final run used the approved agent, cases, tools, data, and environment.
- The agent could not access expected answers or grader logic.
- The run followed security, contract, data, and cost controls.
- The report separates agent failures from service and evaluation failures.
- The report states critical failures, uncertainty, and coverage gaps.
- The required reviewers approved the result.
- The decision, limits, monitoring, and rollback conditions are recorded.
- The evidence package is stored in an approved location.
This SOP uses ideas from these frameworks:
- National Institute of Standards and Technology AI Risk Management Framework
- National Institute of Standards and Technology Generative AI Profile
- ISO/IEC 42001 overview
- Open Worldwide Application Security Project AI Testing Guide
Use the current contract and authoritative source when a framework or requirement applies.