Agent Ops is a file-based task execution and verification workflow:
plan → implement → verify → review → report
This public distribution contains reusable source, safety policy, templates, controlled fixtures, and tests. Mutable approvals, attempts, checkpoints, receipts, logs, benchmark runs, and project-specific operational histories are local-only and excluded from version control.
Version 14.4 adds the separately versioned controlled-engineering-v2 suite:
four harder deterministic fixtures covering recursive multi-file configuration
merging, exact ledger reconciliation, concurrent TTL/LRU caching, and safe ZIP
extraction. The original three-case suite remains unchanged for replication.
Version 14.3.2 gives Codex benchmark executions an equivalent fixed isolation
contract: user configuration and rules are ignored, AGENTS.md loading is
disabled, web search is disabled, strict configuration parsing is enabled, and
spawned commands receive only the core shell environment.
Version 14.3.1 forces Hermes benchmark executions into --safe-mode, excluding
global rules, memory, plugins, MCP customization, and user configuration while
retaining explicit provider/model selection and protected environment
credentials.
Version 14.3 added fixed, disabled-by-default Codex and Hermes benchmark executor
profiles. benchmark execute-case can launch one allowlisted CLI process only
after configuration enablement and explicit --execute, applies an OS-level
hard timeout, runs cases sequentially, redacts and hashes output, and imports
CLI usage when available. It has no arbitrary command or prompt option and
performs no provider call during installation, dry-runs, or tests.
Version 14.0 added evidence-only benchmark baselines. Hashed benchmark runs measure completed-task rate, acceptance-criterion coverage, verification outcomes, overrides, receipt-backed usage, and evidenced cost. Historical cases are never presented as an apples-to-apples model comparison, and unavailable latency, first-attempt, executor, or human-effort data remains unavailable. Benchmark commands do not call models, providers, task verification commands, or application code.
Version 13.1 converts checkpoint filesystem failures into concise CLI errors and generates report headers without trailing whitespace.
Version 13.0 added hashed persistent workflow checkpoints. A clean Git baseline
can be prepared in one session and explicitly resumed in another after an
external build. Resume is dry-run by default and refuses tampering, moved
branches, moved HEAD, or checkpoint reuse. No background process is created.
Version 12.0 added agentops harness-run: a one-iteration, sequential operator
workflow for plan → external build → deterministic verification → review →
summary. It does not edit code, retry, run in parallel, or work in the
background. Dry-run is the default; noninteractive execution requires explicit
external-build and review confirmations. Live summarization still requires its
own approval.
Version 11.0 added agentops summarize: a one-shot, approval-gated OpenRouter
summarizer that activates only in process memory. The provider configuration
stays disabled and dry-run by default. Execution still requires --execute, a
matching approval, credentials, allowlisted routing, spend limits, and all
receipt/attempt integrity controls. An optional hidden terminal prompt keeps a
manually entered key process-only.
Version 10.6 locks the OpenRouter GPT-5 Nano summarizer pilot to minimal reasoning with reasoning text excluded, preserving the tiny output budget for the requested summary while keeping tools and provider fallbacks disabled.
Version 10.5 distinguishes provider network attempts, provider-accepted model requests, and successful verified outputs in reports. OpenRouter validation failures retain only allowlisted response metadata such as finish reason and output/usage presence; raw responses and credentials are never persisted.
Version 10.4 adds a verified macOS CA-bundle fallback for provider HTTPS when the Python installation's configured certificate bundle is missing. Certificate verification remains mandatory; Agent Ops never disables TLS verification.
Version 10.3 adds a credential-free, non-inference OpenRouter connectivity probe. Version 10.2 added read-only Hermes/Codex/OpenRouter integration diagnostics. Version 10.1 added a strictly allowlisted OpenRouter summarizer adapter and a declarative mini-swe-agent harness profile. OpenAI, OpenRouter, and mock adapters remain disabled and dry-run-only by default. The harness is not installed or executed by Agent Ops. No builder, code-editing, tool-use, provider fallback, parallel, or autonomous execution is enabled.
Python 3.10 or newer is recommended.
git clone https://github.com/ondrakingboss/agent-ops.git
cd agent-ops
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txtRun the CLI through the environment:
python agentops --version
python agentops status
python agentops list
python agentops show example-task
python agentops report example-task
python agentops models
python agentops integrations
python agentops probe openrouter
python agentops route lead
python agentops route-task example-task
python agentops invoke example-task --role summarizer \
--provider mock --model mock-summary-v1 \
--input-summary "Summarize the verified task result"
python agentops summarize example-task
python agentops invoke example-task --role summarizer \
--provider openai --model gpt-5-nano \
--input-summary "Summarize the verified task result"
python agentops invoke example-task --role summarizer \
--provider openrouter --model openai/gpt-5-nano \
--input-summary "Summarize the verified task result"
python agentops harnesses
python agentops harness-plan example-task --harness mini-swe-agent
python agentops harness-run example-task
python agentops checkpoint prepare example-task
python agentops checkpoint list example-task
python agentops benchmark list
python agentops benchmark show controlled-engineering-v2
python agentops benchmark verify <benchmark-run-id>
python agentops benchmark report <benchmark-run-id>
python agentops benchmark executors
python agentops receipt list example-task
python agentops receipt usage example-taskRun the regression tests:
python -m unittest discover -s tests -vThe optional live test is skipped unless both explicit opt-in and credentials are present:
AGENTOPS_LIVE_PROVIDER_TEST=1 OPENAI_API_KEY=... \
python -m unittest tests.test_agentops.AgentOpsTests.test_optional_live_openai_summarizer -vThe controlled canary has a separate optional test and also requires a valid approval nonce:
AGENTOPS_LIVE_PROVIDER_TEST=1 OPENAI_API_KEY=... \
python -m unittest tests.test_agentops.AgentOpsTests.test_optional_live_canary -vagentops init
agentops list
agentops show <task-id>
agentops status
agentops run <task-id> [--auto] [--allow-dirty]
[--override-diff] [--override-verification]
agentops report <task-id> [--refresh]
agentops models
agentops integrations
agentops probe openrouter [--timeout-seconds N]
agentops route <role>
agentops route-task <task-id>
agentops harnesses
agentops harness-plan <task-id> [--harness mini-swe-agent]
agentops harness-run <task-id> [--harness agentops-sequential] [--execute]
[--non-interactive --external-build-confirmed
--review-approved]
[--skip-summary | --execute-summary
--summary-approval <nonce> [--prompt-credential]]
[--allow-dirty] [--override-diff]
[--override-verification]
agentops checkpoint prepare <task-id>
agentops checkpoint list [<task-id>]
agentops checkpoint show <checkpoint-id>
agentops checkpoint resume <task-id> --checkpoint <checkpoint-id> [--execute]
[--non-interactive --external-build-confirmed
--review-approved]
[--override-diff] [--override-verification]
agentops benchmark list
agentops benchmark show <suite-id>
agentops benchmark run <suite-id>
agentops benchmark verify <run-id>
agentops benchmark report <run-id>
agentops benchmark executors
agentops benchmark prepare <controlled-suite-id> --executor <label>
[--max-total-tokens N] [--max-api-calls N]
[--max-case-seconds N]
agentops benchmark start <trial-id> [--execute]
agentops benchmark case-start <trial-id> <case-id> [--execute]
agentops benchmark case-finish <trial-id> <case-id>
[--usage-file <cli-usage.json>] [--execute]
agentops benchmark execute-case <trial-id> <case-id>
--profile <fixed-profile> [--execute]
agentops benchmark abandon <trial-id> --reason <reason>
agentops benchmark trial <trial-id>
agentops benchmark evaluate <trial-id> [--execute]
agentops benchmark finalize <trial-id> --human-interventions N --human-minutes N
[--tokens-prompt N] [--tokens-completion N]
[--cost AMOUNT] [--cost-currency USD]
agentops benchmark compare <trial-id> <trial-id> [...]
agentops invoke <task-id> --role <role> --provider <provider> --model <model>
--input-summary <summary> [--execute]
[--approval <nonce>] [--force-new-attempt]
[--max-prompt-tokens N] [--max-completion-tokens N]
[--max-cost AMOUNT] [--timeout-seconds N]
agentops summarize <task-id> [--input-summary <summary>] [--execute]
[--approval <nonce>] [--prompt-credential]
[--max-prompt-tokens N] [--max-completion-tokens N]
[--max-cost AMOUNT] [--timeout-seconds N]
agentops receipt add [<task-id>] [options]
agentops receipt list <task-id>
agentops receipt show <receipt-id>
agentops receipt verify <receipt-id>
agentops receipt revoke <receipt-id> --reason <reason>
agentops receipt resign <receipt-id> --reason <reason>
agentops receipt usage <task-id> [--include-revoked] [--include-superseded]
agentops approve <task-id> --role <role> --provider <provider> --model <model>
--approval-reason <reason> [approval limits]
agentops approve list [<task-id>]
agentops approve show <approval-id>
agentops approve verify <approval-id>
agentops approve revoke <approval-id> --reason <reason>
agentops spend [<task-id>] [--provider openai]
agentops canary openai-summarizer --task-id <task-id>
[--approval <nonce> --execute]
agentops integrations reports the current Hermes route, Codex's
hermes-tools MCP connection, the OpenRouter adapter safety state, and whether
the required credential is present in the current process. It reads only
credential variable names from the Hermes dotenv file, never values, and never
imports Hermes credentials into Agent Ops. The command performs no model or
provider network call.
Provider transport failures are recorded with fixed redacted categories such
as dns, tls_certificate, tls, connection_refused,
connection_interrupted, network_unreachable, timeout, or transport.
Raw socket exception text is never copied into attempts, receipts, or evidence.
agentops probe openrouter performs one GET against the exact allowlisted
public models-catalog endpoint. It sends no API key, prompt, completion request,
tools, or inference payload; reads only the first response byte; and creates no
approval, attempt, receipt, or evidence record.
agentops summarize <task-id> is dry-run by default. It is permanently scoped
to the OpenRouter openai/gpt-5-nano summarizer route with minimal reasoning,
no tools, and no provider fallback. On --execute, a valid matching approval
and credential are required. The adapter becomes active only in that Python
process; config/provider-adapters.yaml is never rewritten. --prompt-credential
uses hidden terminal input and removes the credential from the process after
the command returns.
- --allow-dirty records and permits a pre-existing dirty baseline.
- --override-diff explicitly permits a run without a required Git diff.
- --override-verification explicitly permits completion after failed checks.
- Overrides are persisted in task results, logs, and generated reports.
- Options for run may appear before or after the task ID.
context.project may be an absolute path or a directory below the user's home directory. Agent Ops itself does not need to be stored in a Git repository. Git status and diff evidence are collected from the target project.
If the target project is not a Git repository, Git mode is marked unavailable. A task with requires_diff: true then blocks unless --override-diff is explicitly supplied.
verification:
- name: Backend tests
type: shell
cwd: backend
command: source venv/bin/activate && python -m pytest tests/ -q
timeout: 120
expected_exit: 0
expected_output_contains: passed
expected_output_not_contains: failed
- name: Browser check
type: manual
description: Confirm the changed page renders without console errors.Shell checks run with Bash and pipefail. Full command output is retained in the timestamped log under reports/logs/.
Legacy v1-v5 verification mappings are normalized in memory. Historical task files are not rewritten by read-only commands; a task is upgraded to the current schema only when a write operation updates it.
Benchmark suite manifests live under benchmarks/. The public distribution
ships only synthetic controlled fixtures. Private project histories and their
generated evidence do not belong in a public distribution.
agentops benchmark run reads task, receipt, and attempt records and writes a
hashed YAML snapshot under benchmark-runs/ plus a Markdown report under
reports/benchmarks/. It does not rerun verification or invoke an executor.
Only active, integrity-valid invocation receipts count as model usage.
Historical tasks differ in objective, difficulty, repository state, and execution conditions. Consequently, these reports establish evidence coverage and data-quality baselines; they do not rank models.
Operators may author their own evidence-only suite locally, but should review every referenced task, receipt, attempt, and report before publishing it.
The controlled-engineering suite provides three intentionally broken,
stdlib-only Python fixtures. benchmark prepare copies each fixture into an
ignored disposable directory, initializes an identical Git baseline, hashes
the fixture bundle, and records only an operator-supplied executor label. Agent
Ops does not invoke that executor or any model in the manual lifecycle:
python agentops benchmark prepare controlled-engineering --executor codex
python agentops benchmark start <trial-id> --execute
python agentops benchmark case-start <trial-id> slugify-contract --execute
# The external executor completes TASK.md in the printed workspace.
python agentops benchmark case-finish <trial-id> slugify-contract \
--usage-file /path/to/cli-usage.json --execute
# Repeat case-start/case-finish sequentially for each remaining case.
python agentops benchmark evaluate <trial-id> --execute
python agentops benchmark finalize <trial-id> \
--human-interventions 1 --human-minutes 8
python agentops benchmark compare <trial-a> <trial-b>For more discriminating follow-up trials, use the separately tracked
controlled-engineering-v2 suite. It contains four high-difficulty,
stdlib-only cases with deterministic tests and explicit source-file scopes:
config-layering-contract— recursive merging and validation across two files;ledger-reconciliation-contract— exact Decimal totals and duplicate handling;ttl-cache-contract— thread-safe TTL expiry, LRU eviction, and single-flight factory behavior; andsafe-zip-extract-contract— traversal-resistant extraction with preflight limits and no overwrites.
Use the same lifecycle commands, substituting controlled-engineering-v2 for
the suite ID. Do not mix v1 and v2 trials in a comparison; Agent Ops rejects
cross-suite comparisons.
v14.3 also provides an optional bounded wrapper around two exact CLI shapes. The shipped profiles are disabled and dry-run-only:
python agentops benchmark executors
python agentops benchmark prepare controlled-engineering \
--executor hermes-openrouter-stepfun-3.5-flash
python agentops benchmark start <trial-id> --execute
python agentops benchmark execute-case <trial-id> slugify-contract \
--profile hermes-stepfun-3.5-flashThat final command is record-neutral and invokes nothing. Execution additionally
requires editing the selected profile to enabled: true and dry_run: false,
then adding --execute. The wrapper accepts no arbitrary executable, arguments,
or prompt. It derives a fixed prompt from the tracked case, restricts the case
to its disposable workspace, kills the process group at the lower of the trial
and profile timeout, and runs Hermes in --safe-mode so global customizations
cannot contaminate the result. It redacts output and hashes the local log. Codex
JSONL usage is normalized when present; Hermes usage comes from its explicit usage file.
These are CLI-reported measurements, not independent provider verification.
Mutating lifecycle commands are dry-runs without --execute, except
abandon, which requires an explicit reason. Cases are timed independently and
must run sequentially, so queue time is not counted as executor time.
case-finish may import the allowlisted fields from a JSON CLI usage report;
Agent Ops normalizes and hashes a local evidence copy. This is external-CLI
evidence, not provider verification.
Evaluation runs only the verification command declared by the tracked suite,
preserves its full output in the hashed trial record, requires a change to an
explicitly allowlisted source file, and can happen only once. Unexpected files
fail the case. Python bytecode caches and .DS_Store are recorded separately as
ignored runtime artifacts and are excluded from fixture preparation. A changed
baseline commit, exceeded reconciled limit, or unsuccessful CLI usage report
also fails the case.
The default per-case reconciliation limits are 100,000 total tokens, 10 API
calls, and 300 seconds. Manual trials check these at case-finish; their
external CLI must enforce any hard real-time stop. The bounded wrapper enforces
the time limit during the process and still reconciles token/call limits from
captured usage after exit. Finalization requires explicit human-effort metrics
and uses imported usage/cost evidence when available without allowing operator
values to overwrite it. abandon retains the hashed record and workspaces but
excludes the trial from comparison. Comparison refuses different suites or
fixture hashes and reports descriptive results only.
The policy is stored in config/model-policy.yaml. The models and route commands display the policy; route-task combines it with task type, priority, budget, iteration limit, parallel preference, preferred models, and verification needs.
Routing is recommendation-only for lead, builder, reviewer, summarizer, and reporter. Model routing remains recommendation-only and does not claim those models were called. QA uses invocation_mode: integrated because Agent Ops really runs deterministic shell checks; this is not an LLM invocation. Generated reports always state whether models were actually called.
Receipts are stored as YAML files under receipts/. A manual receipt records an invocation performed outside Agent Ops. It is operator-supplied evidence, not an independently verified provider record. Supplying evidence_path strengthens the record; Agent Ops hashes that evidence file and reports missing or changed evidence.
Interactive entry:
python agentops receipt addNoninteractive entry:
python agentops receipt add example-task \
--non-interactive \
--role lead \
--provider deepseek \
--model "DeepSeek V4 Pro" \
--invocation-mode manual \
--status success \
--input-summary "Planned the benchmark fix" \
--output-summary "Returned a scoped implementation plan"Integrity and reconciliation:
python agentops receipt verify <receipt-id>
python agentops receipt revoke <receipt-id> --reason "Incorrect attribution"
python agentops receipt resign <receipt-id> --reason "Intentional correction"
python agentops receipt usage <task-id>
python agentops report <task-id> --refreshThe canonical SHA-256 receipt hash excludes lifecycle integrity fields such as updated_at, revocation metadata, evidence_state, and superseded_by. Editing invocation content causes verification to fail until receipt resign is run explicitly. Revocation never deletes a receipt. A new receipt can use --supersedes ; Agent Ops records both relationship directions.
Evidence states are self_reported, file_verified, provider_verified, and revoked. Provider-verified evidence may be produced by the local mock adapter or an explicitly enabled OpenAI/OpenRouter summarizer adapter. Mock verification does not prove a model call. Real-provider verification requires a completed response, request ID, exact model identifier, provider-returned token usage, and hash-valid local evidence. OpenRouter additionally records no-fallback and no-tools controls. Duplicate provider/request_id pairs are reported as warnings.
Modes and statuses remain distinct:
- recommendation_only must use skipped and never claims an invocation.
- manual records an operator-reported external invocation.
- integrated is created by an executing adapter; implemented adapters are mock, OpenAI summarizer, and OpenRouter summarizer.
- failed records an unsuccessful attempt.
- fallback_used requires fallback_from and fallback_reason.
The shipped mock declaration in config/provider-adapters.yaml is disabled and dry-run-only. A command without --execute performs a preflight and writes no receipt or evidence:
python agentops invoke example-task \
--role summarizer --provider mock --model mock-summary-v1 \
--input-summary "Summarize the verified benchmark fix"Execution requires an explicit local configuration change to enabled: true and dry_run: false, plus --execute. The adapter then generates a deterministic fake summary, writes local YAML evidence, and creates a hashed provider_verified receipt. Prompt and completion counts are estimates and mock cost is simulated. The receipt stores the supplied summary, not a full prompt. No secrets or API keys are required or logged.
Safety policy is fail-closed for allowed roles/models, prompt and completion token estimates, simulated maximum cost, and timeout.
The pilot uses the allowlisted https://api.openai.com/v1/responses endpoint,
the summarizer role, and gpt-5-nano only. OpenAI documents the Responses API
for direct text generation and identifies GPT-5 nano as a low-cost model suited
to summarization. See the official text-generation guide,
Responses reference,
and model page.
The shipped configuration has enabled: false and dry_run: true. A dry-run
does not require credentials, make a network request, or create evidence:
python agentops invoke example-task \
--role summarizer --provider openai --model gpt-5-nano \
--input-summary "Summarize the verified benchmark fix"A live call requires changing both configuration flags, exporting
OPENAI_API_KEY, creating a valid route-specific approval, and passing both
--approval <nonce> and --execute. The endpoint, role, model, prompt and
completion token ceilings, estimated-cost ceiling, and timeout are all checked
before the request. The request uses store: false, no tools, and a strict
summarization instruction. API keys are never printed or persisted.
Provider-returned request ID, actual model, status, output, and usage are saved
as local evidence. Token usage is authoritative when returned by the provider.
OpenAI does not return currency cost in the response, so cost is calculated
from configured rates and explicitly marked estimated. Review current pricing
before enabling the pilot.
The OpenRouter adapter is limited to the summarizer role and the exact
openai/gpt-5-nano model slug at the allowlisted
https://openrouter.ai/api/v1/chat/completions endpoint. It sends no tools and
sets provider.allow_fallbacks: false, so OpenRouter cannot substitute another
model/provider route.
The shipped configuration has enabled: false and dry_run: true. This command
performs local preflight only and writes no receipt or evidence:
python agentops invoke example-task \
--role summarizer --provider openrouter --model openai/gpt-5-nano \
--input-summary "Summarize the verified benchmark fix"A live call would require explicit configuration enablement, an
OPENROUTER_API_KEY supplied only through the environment, a matching approval,
and --execute. Limits, exact endpoint/model, disabled fallbacks, disabled tools,
timeout, receipt hashing, and evidence hashing are fail-closed. No OpenRouter
call is made by installation or tests. See
runbooks/openrouter-summarizer-pilot.md before considering a supervised pilot.
config/harnesses.yaml declares mini-swe-agent as an external harness candidate.
Agent Ops does not vendor, install, import, or execute it. The shipped profile has
execution_enabled: false, max_iterations: 0, and disables shell execution,
code editing, background execution, parallel execution, and trajectory import.
agentops harnesses displays that policy. agentops harness-plan <task-id> is a
read-only compatibility/safety plan and explicitly reports that no loop was
started. This reuses an existing open-source harness boundary without claiming
autonomy that Agent Ops does not have.
The same manifest declares agentops-sequential, a local one-iteration
operator workflow. agentops harness-run <task-id> is dry-run by default.
--execute delegates task verification and reporting to the existing fail-closed
agentops run implementation. The build remains external: Codex, Hermes, or the
operator makes the code change, while Agent Ops records and verifies it.
Noninteractive execution requires both --external-build-confirmed and
--review-approved; these are explicit operator attestations, not proof that
Agent Ops edited or reviewed code. Summarization is a dry-run unless
--execute-summary and a separate --summary-approval are supplied. There is
no autonomous retry, parallelism, background work, provider fallback, or
builder-model access. See runbooks/bounded-sequential-workflow.md.
agentops checkpoint prepare <task-id> stores a hashed clean Git baseline.
After an external operator or agent makes uncommitted task edits, checkpoint resume can compare those edits to the stored baseline and run the existing
verification/report path. Resume is dry-run unless --execute is supplied and
requires explicit external-build and review confirmations.
Checkpoints are local operational records and are ignored by Git. v13 refuses
non-Git or dirty preparation, tampered records, branch/HEAD changes, and a
second resume. This is persistence of evidence, not persistence of a running
agent. See runbooks/persistent-checkpoints.md.
Approvals are local SHA-256 integrity records under approvals/. They are
short-lived and bind one task, role, provider, and model to maximum calls,
prompt/completion tokens, and total cost. The local operator is recorded as the
approver. Revocation and expiry are fail-closed and non-destructive. The hash
detects record edits; it is not a remote identity signature.
Every real --execute submission creates a hashed record under attempts/.
The invocation fingerprint covers task, route, normalized input summary, and
approval ID. An in-flight or successful matching fingerprint blocks replay.
--force-new-attempt is explicit, remains subject to every approval and spend
limit, and links the override to the previous attempt.
Spend accounting is local-only and uses active, non-revoked, non-superseded receipts. Configured controls include per-task, provider-daily, provider-monthly, and live-call ceilings. Blocked estimated costs are displayed separately and never represented as provider billing.
See runbooks/openai-summarizer-canary.md for the gated first-call procedure.
agent-ops/
├── agentops
├── agentops_lib/
│ ├── cli.py
│ ├── common.py
│ ├── git_state.py
│ ├── harnesses.py
│ ├── integrations.py
│ ├── probes.py
│ ├── openai_provider.py
│ ├── openrouter_provider.py
│ ├── approvals.py
│ ├── attempts.py
│ ├── benchmarks.py
│ ├── benchmark_executors.py
│ ├── benchmark_trials.py
│ ├── checkpoints.py
│ ├── operations.py
│ ├── orchestration.py
│ ├── providers.py
│ ├── receipts.py
│ ├── reports.py
│ ├── routing.py
│ ├── summaries.py
│ ├── tasks.py
│ ├── verification.py
│ └── workflow.py
├── requirements.txt
├── config/
│ ├── benchmark-executors.yaml
│ ├── harnesses.yaml
│ └── provider-adapters.yaml
├── prompts/
├── approvals/
├── attempts/
├── benchmarks/
│ └── fixtures/
├── benchmark-runs/
├── benchmark-trials/
├── checkpoints/
├── receipts/
│ └── evidence/
├── runbooks/
├── tasks/
│ └── example-task.yaml
├── templates/
├── reports/
│ ├── example-task.md
│ └── logs/
└── tests/
Providers and bounded benchmark executors ship disabled and dry-run-only.
Never commit .env files, API keys, mutable evidence records, or raw prompts.
See SECURITY.md for reporting vulnerabilities and the supported
security boundary.
Contributions are welcome under CONTRIBUTING.md. Agent Ops is released under the MIT License.
- Required diffs are never overridden implicitly.
- Failed verification blocks completion unless explicitly overridden.
- Git failures are not interpreted as dirty files.
- Acceptance criteria require direct linked evidence to be marked verified.
- Reports distinguish failures, unverified work, overrides, and non-Git limits.
- Worker output remains short, but command logs preserve complete evidence.