Skip to content
ondrakingbossPublic

About

Fail-closed, evidence-first workflow for AI-assisted engineering

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Agent Ops — Lightweight AI Agency Workflow

Agent Ops is a file-based task execution and verification workflow:

plan → implement → verify → review → report

CI

This public distribution contains reusable source, safety policy, templates, controlled fixtures, and tests. Mutable approvals, attempts, checkpoints, receipts, logs, benchmark runs, and project-specific operational histories are local-only and excluded from version control.

Version 14.4 adds the separately versioned controlled-engineering-v2 suite: four harder deterministic fixtures covering recursive multi-file configuration merging, exact ledger reconciliation, concurrent TTL/LRU caching, and safe ZIP extraction. The original three-case suite remains unchanged for replication.

Version 14.3.2 gives Codex benchmark executions an equivalent fixed isolation contract: user configuration and rules are ignored, AGENTS.md loading is disabled, web search is disabled, strict configuration parsing is enabled, and spawned commands receive only the core shell environment.

Version 14.3.1 forces Hermes benchmark executions into --safe-mode, excluding global rules, memory, plugins, MCP customization, and user configuration while retaining explicit provider/model selection and protected environment credentials.

Version 14.3 added fixed, disabled-by-default Codex and Hermes benchmark executor profiles. benchmark execute-case can launch one allowlisted CLI process only after configuration enablement and explicit --execute, applies an OS-level hard timeout, runs cases sequentially, redacts and hashes output, and imports CLI usage when available. It has no arbitrary command or prompt option and performs no provider call during installation, dry-runs, or tests.

Version 14.0 added evidence-only benchmark baselines. Hashed benchmark runs measure completed-task rate, acceptance-criterion coverage, verification outcomes, overrides, receipt-backed usage, and evidenced cost. Historical cases are never presented as an apples-to-apples model comparison, and unavailable latency, first-attempt, executor, or human-effort data remains unavailable. Benchmark commands do not call models, providers, task verification commands, or application code.

Version 13.1 converts checkpoint filesystem failures into concise CLI errors and generates report headers without trailing whitespace.

Version 13.0 added hashed persistent workflow checkpoints. A clean Git baseline can be prepared in one session and explicitly resumed in another after an external build. Resume is dry-run by default and refuses tampering, moved branches, moved HEAD, or checkpoint reuse. No background process is created.

Version 12.0 added agentops harness-run: a one-iteration, sequential operator workflow for plan → external build → deterministic verification → review → summary. It does not edit code, retry, run in parallel, or work in the background. Dry-run is the default; noninteractive execution requires explicit external-build and review confirmations. Live summarization still requires its own approval.

Version 11.0 added agentops summarize: a one-shot, approval-gated OpenRouter summarizer that activates only in process memory. The provider configuration stays disabled and dry-run by default. Execution still requires --execute, a matching approval, credentials, allowlisted routing, spend limits, and all receipt/attempt integrity controls. An optional hidden terminal prompt keeps a manually entered key process-only.

Version 10.6 locks the OpenRouter GPT-5 Nano summarizer pilot to minimal reasoning with reasoning text excluded, preserving the tiny output budget for the requested summary while keeping tools and provider fallbacks disabled.

Version 10.5 distinguishes provider network attempts, provider-accepted model requests, and successful verified outputs in reports. OpenRouter validation failures retain only allowlisted response metadata such as finish reason and output/usage presence; raw responses and credentials are never persisted.

Version 10.4 adds a verified macOS CA-bundle fallback for provider HTTPS when the Python installation's configured certificate bundle is missing. Certificate verification remains mandatory; Agent Ops never disables TLS verification.

Version 10.3 adds a credential-free, non-inference OpenRouter connectivity probe. Version 10.2 added read-only Hermes/Codex/OpenRouter integration diagnostics. Version 10.1 added a strictly allowlisted OpenRouter summarizer adapter and a declarative mini-swe-agent harness profile. OpenAI, OpenRouter, and mock adapters remain disabled and dry-run-only by default. The harness is not installed or executed by Agent Ops. No builder, code-editing, tool-use, provider fallback, parallel, or autonomous execution is enabled.

Install

Python 3.10 or newer is recommended.

git clone https://github.com/ondrakingboss/agent-ops.git
cd agent-ops
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt

Run the CLI through the environment:

python agentops --version
python agentops status
python agentops list
python agentops show example-task
python agentops report example-task
python agentops models
python agentops integrations
python agentops probe openrouter
python agentops route lead
python agentops route-task example-task
python agentops invoke example-task --role summarizer \
  --provider mock --model mock-summary-v1 \
  --input-summary "Summarize the verified task result"
python agentops summarize example-task
python agentops invoke example-task --role summarizer \
  --provider openai --model gpt-5-nano \
  --input-summary "Summarize the verified task result"
python agentops invoke example-task --role summarizer \
  --provider openrouter --model openai/gpt-5-nano \
  --input-summary "Summarize the verified task result"
python agentops harnesses
python agentops harness-plan example-task --harness mini-swe-agent
python agentops harness-run example-task
python agentops checkpoint prepare example-task
python agentops checkpoint list example-task
python agentops benchmark list
python agentops benchmark show controlled-engineering-v2
python agentops benchmark verify <benchmark-run-id>
python agentops benchmark report <benchmark-run-id>
python agentops benchmark executors
python agentops receipt list example-task
python agentops receipt usage example-task

Run the regression tests:

python -m unittest discover -s tests -v

The optional live test is skipped unless both explicit opt-in and credentials are present:

AGENTOPS_LIVE_PROVIDER_TEST=1 OPENAI_API_KEY=... \
  python -m unittest tests.test_agentops.AgentOpsTests.test_optional_live_openai_summarizer -v

The controlled canary has a separate optional test and also requires a valid approval nonce:

AGENTOPS_LIVE_PROVIDER_TEST=1 OPENAI_API_KEY=... \
  python -m unittest tests.test_agentops.AgentOpsTests.test_optional_live_canary -v

Commands

agentops init
agentops list
agentops show <task-id>
agentops status
agentops run <task-id> [--auto] [--allow-dirty]
                       [--override-diff] [--override-verification]
agentops report <task-id> [--refresh]
agentops models
agentops integrations
agentops probe openrouter [--timeout-seconds N]
agentops route <role>
agentops route-task <task-id>
agentops harnesses
agentops harness-plan <task-id> [--harness mini-swe-agent]
agentops harness-run <task-id> [--harness agentops-sequential] [--execute]
                     [--non-interactive --external-build-confirmed
                      --review-approved]
                     [--skip-summary | --execute-summary
                      --summary-approval <nonce> [--prompt-credential]]
                     [--allow-dirty] [--override-diff]
                     [--override-verification]
agentops checkpoint prepare <task-id>
agentops checkpoint list [<task-id>]
agentops checkpoint show <checkpoint-id>
agentops checkpoint resume <task-id> --checkpoint <checkpoint-id> [--execute]
                    [--non-interactive --external-build-confirmed
                     --review-approved]
                    [--override-diff] [--override-verification]
agentops benchmark list
agentops benchmark show <suite-id>
agentops benchmark run <suite-id>
agentops benchmark verify <run-id>
agentops benchmark report <run-id>
agentops benchmark executors
agentops benchmark prepare <controlled-suite-id> --executor <label>
                   [--max-total-tokens N] [--max-api-calls N]
                   [--max-case-seconds N]
agentops benchmark start <trial-id> [--execute]
agentops benchmark case-start <trial-id> <case-id> [--execute]
agentops benchmark case-finish <trial-id> <case-id>
                   [--usage-file <cli-usage.json>] [--execute]
agentops benchmark execute-case <trial-id> <case-id>
                   --profile <fixed-profile> [--execute]
agentops benchmark abandon <trial-id> --reason <reason>
agentops benchmark trial <trial-id>
agentops benchmark evaluate <trial-id> [--execute]
agentops benchmark finalize <trial-id> --human-interventions N --human-minutes N
                   [--tokens-prompt N] [--tokens-completion N]
                   [--cost AMOUNT] [--cost-currency USD]
agentops benchmark compare <trial-id> <trial-id> [...]
agentops invoke <task-id> --role <role> --provider <provider> --model <model>
                --input-summary <summary> [--execute]
                [--approval <nonce>] [--force-new-attempt]
                [--max-prompt-tokens N] [--max-completion-tokens N]
                [--max-cost AMOUNT] [--timeout-seconds N]
agentops summarize <task-id> [--input-summary <summary>] [--execute]
                   [--approval <nonce>] [--prompt-credential]
                   [--max-prompt-tokens N] [--max-completion-tokens N]
                   [--max-cost AMOUNT] [--timeout-seconds N]
agentops receipt add [<task-id>] [options]
agentops receipt list <task-id>
agentops receipt show <receipt-id>
agentops receipt verify <receipt-id>
agentops receipt revoke <receipt-id> --reason <reason>
agentops receipt resign <receipt-id> --reason <reason>
agentops receipt usage <task-id> [--include-revoked] [--include-superseded]
agentops approve <task-id> --role <role> --provider <provider> --model <model>
                 --approval-reason <reason> [approval limits]
agentops approve list [<task-id>]
agentops approve show <approval-id>
agentops approve verify <approval-id>
agentops approve revoke <approval-id> --reason <reason>
agentops spend [<task-id>] [--provider openai]
agentops canary openai-summarizer --task-id <task-id>
                [--approval <nonce> --execute]

agentops integrations reports the current Hermes route, Codex's hermes-tools MCP connection, the OpenRouter adapter safety state, and whether the required credential is present in the current process. It reads only credential variable names from the Hermes dotenv file, never values, and never imports Hermes credentials into Agent Ops. The command performs no model or provider network call.

Provider transport failures are recorded with fixed redacted categories such as dns, tls_certificate, tls, connection_refused, connection_interrupted, network_unreachable, timeout, or transport. Raw socket exception text is never copied into attempts, receipts, or evidence.

agentops probe openrouter performs one GET against the exact allowlisted public models-catalog endpoint. It sends no API key, prompt, completion request, tools, or inference payload; reads only the first response byte; and creates no approval, attempt, receipt, or evidence record.

agentops summarize <task-id> is dry-run by default. It is permanently scoped to the OpenRouter openai/gpt-5-nano summarizer route with minimal reasoning, no tools, and no provider fallback. On --execute, a valid matching approval and credential are required. The adapter becomes active only in that Python process; config/provider-adapters.yaml is never rewritten. --prompt-credential uses hidden terminal input and removes the credential from the process after the command returns.

  • --allow-dirty records and permits a pre-existing dirty baseline.
  • --override-diff explicitly permits a run without a required Git diff.
  • --override-verification explicitly permits completion after failed checks.
  • Overrides are persisted in task results, logs, and generated reports.
  • Options for run may appear before or after the task ID.

Project resolution

context.project may be an absolute path or a directory below the user's home directory. Agent Ops itself does not need to be stored in a Git repository. Git status and diff evidence are collected from the target project.

If the target project is not a Git repository, Git mode is marked unavailable. A task with requires_diff: true then blocks unless --override-diff is explicitly supplied.

Structured verification

verification:
  - name: Backend tests
    type: shell
    cwd: backend
    command: source venv/bin/activate && python -m pytest tests/ -q
    timeout: 120
    expected_exit: 0
    expected_output_contains: passed
    expected_output_not_contains: failed
  - name: Browser check
    type: manual
    description: Confirm the changed page renders without console errors.

Shell checks run with Bash and pipefail. Full command output is retained in the timestamped log under reports/logs/.

Legacy v1-v5 verification mappings are normalized in memory. Historical task files are not rewritten by read-only commands; a task is upgraded to the current schema only when a write operation updates it.

Evidence-only benchmarks

Benchmark suite manifests live under benchmarks/. The public distribution ships only synthetic controlled fixtures. Private project histories and their generated evidence do not belong in a public distribution.

agentops benchmark run reads task, receipt, and attempt records and writes a hashed YAML snapshot under benchmark-runs/ plus a Markdown report under reports/benchmarks/. It does not rerun verification or invoke an executor. Only active, integrity-valid invocation receipts count as model usage.

Historical tasks differ in objective, difficulty, repository state, and execution conditions. Consequently, these reports establish evidence coverage and data-quality baselines; they do not rank models.

Operators may author their own evidence-only suite locally, but should review every referenced task, receipt, attempt, and report before publishing it.

The controlled-engineering suite provides three intentionally broken, stdlib-only Python fixtures. benchmark prepare copies each fixture into an ignored disposable directory, initializes an identical Git baseline, hashes the fixture bundle, and records only an operator-supplied executor label. Agent Ops does not invoke that executor or any model in the manual lifecycle:

python agentops benchmark prepare controlled-engineering --executor codex
python agentops benchmark start <trial-id> --execute
python agentops benchmark case-start <trial-id> slugify-contract --execute
# The external executor completes TASK.md in the printed workspace.
python agentops benchmark case-finish <trial-id> slugify-contract \
  --usage-file /path/to/cli-usage.json --execute
# Repeat case-start/case-finish sequentially for each remaining case.
python agentops benchmark evaluate <trial-id> --execute
python agentops benchmark finalize <trial-id> \
  --human-interventions 1 --human-minutes 8
python agentops benchmark compare <trial-a> <trial-b>

For more discriminating follow-up trials, use the separately tracked controlled-engineering-v2 suite. It contains four high-difficulty, stdlib-only cases with deterministic tests and explicit source-file scopes:

  • config-layering-contract — recursive merging and validation across two files;
  • ledger-reconciliation-contract — exact Decimal totals and duplicate handling;
  • ttl-cache-contract — thread-safe TTL expiry, LRU eviction, and single-flight factory behavior; and
  • safe-zip-extract-contract — traversal-resistant extraction with preflight limits and no overwrites.

Use the same lifecycle commands, substituting controlled-engineering-v2 for the suite ID. Do not mix v1 and v2 trials in a comparison; Agent Ops rejects cross-suite comparisons.

v14.3 also provides an optional bounded wrapper around two exact CLI shapes. The shipped profiles are disabled and dry-run-only:

python agentops benchmark executors
python agentops benchmark prepare controlled-engineering \
  --executor hermes-openrouter-stepfun-3.5-flash
python agentops benchmark start <trial-id> --execute
python agentops benchmark execute-case <trial-id> slugify-contract \
  --profile hermes-stepfun-3.5-flash

That final command is record-neutral and invokes nothing. Execution additionally requires editing the selected profile to enabled: true and dry_run: false, then adding --execute. The wrapper accepts no arbitrary executable, arguments, or prompt. It derives a fixed prompt from the tracked case, restricts the case to its disposable workspace, kills the process group at the lower of the trial and profile timeout, and runs Hermes in --safe-mode so global customizations cannot contaminate the result. It redacts output and hashes the local log. Codex JSONL usage is normalized when present; Hermes usage comes from its explicit usage file. These are CLI-reported measurements, not independent provider verification.

Mutating lifecycle commands are dry-runs without --execute, except abandon, which requires an explicit reason. Cases are timed independently and must run sequentially, so queue time is not counted as executor time. case-finish may import the allowlisted fields from a JSON CLI usage report; Agent Ops normalizes and hashes a local evidence copy. This is external-CLI evidence, not provider verification.

Evaluation runs only the verification command declared by the tracked suite, preserves its full output in the hashed trial record, requires a change to an explicitly allowlisted source file, and can happen only once. Unexpected files fail the case. Python bytecode caches and .DS_Store are recorded separately as ignored runtime artifacts and are excluded from fixture preparation. A changed baseline commit, exceeded reconciled limit, or unsuccessful CLI usage report also fails the case.

The default per-case reconciliation limits are 100,000 total tokens, 10 API calls, and 300 seconds. Manual trials check these at case-finish; their external CLI must enforce any hard real-time stop. The bounded wrapper enforces the time limit during the process and still reconciles token/call limits from captured usage after exit. Finalization requires explicit human-effort metrics and uses imported usage/cost evidence when available without allowing operator values to overwrite it. abandon retains the hashed record and workspaces but excludes the trial from comparison. Comparison refuses different suites or fixture hashes and reports descriptive results only.

Model routing

The policy is stored in config/model-policy.yaml. The models and route commands display the policy; route-task combines it with task type, priority, budget, iteration limit, parallel preference, preferred models, and verification needs.

Routing is recommendation-only for lead, builder, reviewer, summarizer, and reporter. Model routing remains recommendation-only and does not claim those models were called. QA uses invocation_mode: integrated because Agent Ops really runs deterministic shell checks; this is not an LLM invocation. Generated reports always state whether models were actually called.

Invocation receipts

Receipts are stored as YAML files under receipts/. A manual receipt records an invocation performed outside Agent Ops. It is operator-supplied evidence, not an independently verified provider record. Supplying evidence_path strengthens the record; Agent Ops hashes that evidence file and reports missing or changed evidence.

Interactive entry:

python agentops receipt add

Noninteractive entry:

python agentops receipt add example-task \
  --non-interactive \
  --role lead \
  --provider deepseek \
  --model "DeepSeek V4 Pro" \
  --invocation-mode manual \
  --status success \
  --input-summary "Planned the benchmark fix" \
  --output-summary "Returned a scoped implementation plan"

Integrity and reconciliation:

python agentops receipt verify <receipt-id>
python agentops receipt revoke <receipt-id> --reason "Incorrect attribution"
python agentops receipt resign <receipt-id> --reason "Intentional correction"
python agentops receipt usage <task-id>
python agentops report <task-id> --refresh

The canonical SHA-256 receipt hash excludes lifecycle integrity fields such as updated_at, revocation metadata, evidence_state, and superseded_by. Editing invocation content causes verification to fail until receipt resign is run explicitly. Revocation never deletes a receipt. A new receipt can use --supersedes ; Agent Ops records both relationship directions.

Evidence states are self_reported, file_verified, provider_verified, and revoked. Provider-verified evidence may be produced by the local mock adapter or an explicitly enabled OpenAI/OpenRouter summarizer adapter. Mock verification does not prove a model call. Real-provider verification requires a completed response, request ID, exact model identifier, provider-returned token usage, and hash-valid local evidence. OpenRouter additionally records no-fallback and no-tools controls. Duplicate provider/request_id pairs are reported as warnings.

Modes and statuses remain distinct:

  • recommendation_only must use skipped and never claims an invocation.
  • manual records an operator-reported external invocation.
  • integrated is created by an executing adapter; implemented adapters are mock, OpenAI summarizer, and OpenRouter summarizer.
  • failed records an unsuccessful attempt.
  • fallback_used requires fallback_from and fallback_reason.

Mock invocation

The shipped mock declaration in config/provider-adapters.yaml is disabled and dry-run-only. A command without --execute performs a preflight and writes no receipt or evidence:

python agentops invoke example-task \
  --role summarizer --provider mock --model mock-summary-v1 \
  --input-summary "Summarize the verified benchmark fix"

Execution requires an explicit local configuration change to enabled: true and dry_run: false, plus --execute. The adapter then generates a deterministic fake summary, writes local YAML evidence, and creates a hashed provider_verified receipt. Prompt and completion counts are estimates and mock cost is simulated. The receipt stores the supplied summary, not a full prompt. No secrets or API keys are required or logged.

Safety policy is fail-closed for allowed roles/models, prompt and completion token estimates, simulated maximum cost, and timeout.

OpenAI summarizer pilot

The pilot uses the allowlisted https://api.openai.com/v1/responses endpoint, the summarizer role, and gpt-5-nano only. OpenAI documents the Responses API for direct text generation and identifies GPT-5 nano as a low-cost model suited to summarization. See the official text-generation guide, Responses reference, and model page.

The shipped configuration has enabled: false and dry_run: true. A dry-run does not require credentials, make a network request, or create evidence:

python agentops invoke example-task \
  --role summarizer --provider openai --model gpt-5-nano \
  --input-summary "Summarize the verified benchmark fix"

A live call requires changing both configuration flags, exporting OPENAI_API_KEY, creating a valid route-specific approval, and passing both --approval <nonce> and --execute. The endpoint, role, model, prompt and completion token ceilings, estimated-cost ceiling, and timeout are all checked before the request. The request uses store: false, no tools, and a strict summarization instruction. API keys are never printed or persisted.

Provider-returned request ID, actual model, status, output, and usage are saved as local evidence. Token usage is authoritative when returned by the provider. OpenAI does not return currency cost in the response, so cost is calculated from configured rates and explicitly marked estimated. Review current pricing before enabling the pilot.

OpenRouter summarizer pilot

The OpenRouter adapter is limited to the summarizer role and the exact openai/gpt-5-nano model slug at the allowlisted https://openrouter.ai/api/v1/chat/completions endpoint. It sends no tools and sets provider.allow_fallbacks: false, so OpenRouter cannot substitute another model/provider route.

The shipped configuration has enabled: false and dry_run: true. This command performs local preflight only and writes no receipt or evidence:

python agentops invoke example-task \
  --role summarizer --provider openrouter --model openai/gpt-5-nano \
  --input-summary "Summarize the verified benchmark fix"

A live call would require explicit configuration enablement, an OPENROUTER_API_KEY supplied only through the environment, a matching approval, and --execute. Limits, exact endpoint/model, disabled fallbacks, disabled tools, timeout, receipt hashing, and evidence hashing are fail-closed. No OpenRouter call is made by installation or tests. See runbooks/openrouter-summarizer-pilot.md before considering a supervised pilot.

Bounded sequential harness

config/harnesses.yaml declares mini-swe-agent as an external harness candidate. Agent Ops does not vendor, install, import, or execute it. The shipped profile has execution_enabled: false, max_iterations: 0, and disables shell execution, code editing, background execution, parallel execution, and trajectory import.

agentops harnesses displays that policy. agentops harness-plan <task-id> is a read-only compatibility/safety plan and explicitly reports that no loop was started. This reuses an existing open-source harness boundary without claiming autonomy that Agent Ops does not have.

The same manifest declares agentops-sequential, a local one-iteration operator workflow. agentops harness-run <task-id> is dry-run by default. --execute delegates task verification and reporting to the existing fail-closed agentops run implementation. The build remains external: Codex, Hermes, or the operator makes the code change, while Agent Ops records and verifies it.

Noninteractive execution requires both --external-build-confirmed and --review-approved; these are explicit operator attestations, not proof that Agent Ops edited or reviewed code. Summarization is a dry-run unless --execute-summary and a separate --summary-approval are supplied. There is no autonomous retry, parallelism, background work, provider fallback, or builder-model access. See runbooks/bounded-sequential-workflow.md.

Persistent checkpoints

agentops checkpoint prepare <task-id> stores a hashed clean Git baseline. After an external operator or agent makes uncommitted task edits, checkpoint resume can compare those edits to the stored baseline and run the existing verification/report path. Resume is dry-run unless --execute is supplied and requires explicit external-build and review confirmations.

Checkpoints are local operational records and are ignored by Git. v13 refuses non-Git or dirty preparation, tampered records, branch/HEAD changes, and a second resume. This is persistence of evidence, not persistence of a running agent. See runbooks/persistent-checkpoints.md.

Live pilot operational controls

Approvals are local SHA-256 integrity records under approvals/. They are short-lived and bind one task, role, provider, and model to maximum calls, prompt/completion tokens, and total cost. The local operator is recorded as the approver. Revocation and expiry are fail-closed and non-destructive. The hash detects record edits; it is not a remote identity signature.

Every real --execute submission creates a hashed record under attempts/. The invocation fingerprint covers task, route, normalized input summary, and approval ID. An in-flight or successful matching fingerprint blocks replay. --force-new-attempt is explicit, remains subject to every approval and spend limit, and links the override to the previous attempt.

Spend accounting is local-only and uses active, non-revoked, non-superseded receipts. Configured controls include per-task, provider-daily, provider-monthly, and live-call ceilings. Blocked estimated costs are displayed separately and never represented as provider billing.

See runbooks/openai-summarizer-canary.md for the gated first-call procedure.

Directory structure

agent-ops/
├── agentops
├── agentops_lib/
│   ├── cli.py
│   ├── common.py
│   ├── git_state.py
│   ├── harnesses.py
│   ├── integrations.py
│   ├── probes.py
│   ├── openai_provider.py
│   ├── openrouter_provider.py
│   ├── approvals.py
│   ├── attempts.py
│   ├── benchmarks.py
│   ├── benchmark_executors.py
│   ├── benchmark_trials.py
│   ├── checkpoints.py
│   ├── operations.py
│   ├── orchestration.py
│   ├── providers.py
│   ├── receipts.py
│   ├── reports.py
│   ├── routing.py
│   ├── summaries.py
│   ├── tasks.py
│   ├── verification.py
│   └── workflow.py
├── requirements.txt
├── config/
│   ├── benchmark-executors.yaml
│   ├── harnesses.yaml
│   └── provider-adapters.yaml
├── prompts/
├── approvals/
├── attempts/
├── benchmarks/
│   └── fixtures/
├── benchmark-runs/
├── benchmark-trials/
├── checkpoints/
├── receipts/
│   └── evidence/
├── runbooks/
├── tasks/
│   └── example-task.yaml
├── templates/
├── reports/
│   ├── example-task.md
│   └── logs/
└── tests/

Security and disclosure

Providers and bounded benchmark executors ship disabled and dry-run-only. Never commit .env files, API keys, mutable evidence records, or raw prompts. See SECURITY.md for reporting vulnerabilities and the supported security boundary.

Contributing and license

Contributions are welcome under CONTRIBUTING.md. Agent Ops is released under the MIT License.

Reliability rules

  1. Required diffs are never overridden implicitly.
  2. Failed verification blocks completion unless explicitly overridden.
  3. Git failures are not interpreted as dirty files.
  4. Acceptance criteria require direct linked evidence to be marked verified.
  5. Reports distinguish failures, unverified work, overrides, and non-Git limits.
  6. Worker output remains short, but command logs preserve complete evidence.

About

Fail-closed, evidence-first workflow for AI-assisted engineering

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages