Skip to content

Latest commit

 

History

142 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Token Economy

Token economics for LLM coding agents — one tested source for model pricing (with history), run-cost computation, and token-efficiency model selection. .NET, dependency-free core.

NuGet NuGet downloads CI License: Apache 2.0

dotnet add package TokenEconomy

Package page: TokenEconomy · GitHub releases

An agent run bills for input, output, and cache tokens at a rate that changes over time, per model. Getting that wrong is not a rounding error: a hard-coded price silently costs the wrong amount for every historic run, and a missing price that defaults to 0 reports a budget as healthy while it drains. All 25 catalog models have dated Standard API price histories and primary-source provenance, verified through 2026-09-25. See price history and scope for historical corrections, cache semantics, and API cost versus subscription consumption.

Direct code-review capability has its own benchmark category: 24 sourced measurements across five studies, with precision, known-issue coverage and original review configurations. See the review research and C# guide.

TokenEconomy keeps the prices — with their validity dates — and answers two questions from that one source: what did this run cost? and which model buys the most for the tokens I have left? Unknown is always returned as unknown.

API guide: https://agent-orchestrator.dev/token-economy/api/ — inputs, return types, reasoning effort, scores, and executable C# examples.

Docs & website: https://agent-orchestrator.dev/token-economy/ — a static site built from website/ and deployed by deploy-website.yml (see website/DEPLOY.md).

Research plan: Forecast each task as a percentage of a five-hour cap defines the proposed measurement, uncertainty, repository boundaries, and GO-blocked delivery slices. Its visual explainers are part of the static site. No forecast implementation is authorised yet.

Delegation economy: the pattern guide explains how orchestrators assign bounded work to cheaper model tiers and when to escalate. Prompts and task cards can include the standardized context block verbatim. Task-cutting guide: Dynamic workflows as a task-cutting strategy compares one Claude workflow-sized card with small dependsOn cards using Agent Studio token, review, retry, and gate evidence. It includes the current Claude-only workflow constraint, Codex integration gap, TE-8 routing hook, and a reusable AI-pattern candidate.

Prompt-enrichment analysis: Prompt enrichment before an agent run quantifies rule, embedding, classifier, and hybrid preprocessing against the Agent Studio retry baseline. It specifies the auditable enrichment-report.json contract, a Task Server integration boundary, a selection rubric, and a standalone ai-patterns handoff.

What it does

  • Pricing catalog with history — per-model price entries keyed by ValidFrom; a run's cost is computed with the prices valid at the run's timestamp. Historic entries are kept, never overwritten. Every listing also carries a canonical DisplayName, so a consumer never has to maintain its own partial, drifting model-label table.

  • Cost API — ComputeCost(model, usage, atUtc) → deterministic breakdown + total; unknown models return an explicit unknown, never a silent zero.

  • Versioned model-routing policy — one embedded, schema-backed knowledge base resolves model aliases, providers/CLIs, reasoning levels, score tiers, correctness floors, workflow exceptions, restrictions, deprecations, reissues, and evidence status. The core ladder follows Agent Studio; the local policy revision adds explicit Astra, Fable 5.1, Claude Opus 5.5, GPT-6 Sol, and GPT-6 Luna support. The versioned policy is authoritative; pricing and quota cannot lower its correctness floors. See the generated knowledge view and authoritative policy.

  • Token-efficiency compatibility matrix — SuggestModel(taskClass, budgetPressure, availableClis, atUtc) ranks only policy-selectable core models. Use ModelRoutingPolicy for correctness decisions and ModelRoutingKnowledgeBase.FallbacksFor(routeId) for evidence-scoped provider fallbacks. Cost class is derived from the pricing catalog, never restated.

  • Agent Studio ingest — reads each card's durable task.json directly and upserts model-run metrics by task key + run. It maps model/provider, usage, price-at-run-time, final-lane outcome signal, timestamps, project, task type, CLI and thinking level; ModelRunViews provides daily per-model, CLI, and project consumption/outcome views with explicit unresolved-cost counts. The filesystem contract is intentionally used over the task-server API so reporting jobs do not require a running server.

  • Provider availability snapshot — captures provider/CLI probe state, independently identified quota windows, observed usage/headroom and reset, freshness, conservative warning state, and model-price coverage at one decision instant. Imported runs supply only the separately labelled rate and exhaustion projection; they never replace observed quota. Unknown, stale, suspicious, unavailable, or unpriced inputs cannot render as healthy or as a zero cost. The snapshot is routing evidence, not model selection. See the data contract and rendered snapshot.

  • Deterministic routing composition — ModelRouter.Route combines an upfront estimate, the versioned policy and knowledge base, benchmark and trust qualification, workflow capabilities, available CLIs, run-scoped quota/budget state, and an optional operator pin. It returns a selected model/thinking level or an explicit wait/override decision without allowing quota or cost to cross a correctness floor. See the routing API contract.

  • Agent Studio admission loop — AgentStudioTaskAdmission.PrepareAttempt stores intake features, consumes the newest classified prior outcome, routes before each attempt against a run-scoped quota snapshot, records the complete decision, and returns either an attempt-local launch route or a closed wait/override disposition. Card configuration is retained separately and is never silently rewritten. See the host integration contract.

  • Completed-card economics: compare model/reasoning pairs using local retries, dated token costs, review evidence, organisation refusals, and subscription quota share through the decision API and CLI.

  • Controlled A/B benchmarks — versioned repository definitions execute the same task against model/effort variants in isolated workspaces, retain raw append-only measurements, and derive deterministic comparison reports. Start with the end-to-end benchmark methodology; the protocol background records the established-suite influences.

  • External benchmark evidence and price-performance — schema-backed, append-only benchmark definitions and model/effort measurements share one typed catalog with controlled setups. ModelBenchmarkMatrix returns raw or weighted score cells, dated token prices, published task metrics, declared- assumption cost, score per dollar, reference deltas, evidence age, and ranked cheaper-and-at-least-as-good candidates. It recommends; it does not route or weaken policy floors. See the public matrix and refresh recipe.

  • Human-friendly language capability — dated German/English evidence for six language dimensions, typed overall and per-dimension routing constraints, and sample-cost ranking with stable ties. A validated Voice Lint importer retains source hashes and history; the capability guide and website matrix distinguish public research from measured model scores.

  • Study-backed task-class advice — TaskClassRecommendationCatalog.Default.Recommend(taskClass) returns the ranked equivalent model/thinking set, rationale version, evidence, measured cost per outcome, and downgrade boundary used at card creation; a separate Select(set, quotaState, atUtc) chooses concretely only from fresh quota evidence. OutcomeEfficiency.ComputeObserved calculates tokens per accepted outcome and retry-adjusted dated list-price cost. The published task-class studies cover the complete taxonomy and the first controlled HTML/UI and source-review pilots; the concrete card still goes through ModelRouter so correctness floors win.

  • Upfront task complexity — card/repository signals plus measured historical neighbours produce a versioned routing score, confidence, token/reissue forecast, and audit evidence. A host-supplied mini-model rubric is optional. See the design and backtest contract.

  • Document-to-text benchmark — a curated PDF/Word/spreadsheet/presentation hard-case corpus runs across every catalog model and derives evidence-linked, per-document-type capability records without turning failures into unsupported claims. See the capability-corpus methodology.

  • Native media capability catalog — evidence-dated image, video, music, speech, and dictation rows for Codex, Antigravity, and Claude Code are pulled from the same embedded catalog convention as pricing. Includes the retained N=4 Codex image benchmark and explicit unknown/unverified cost factors. See the capability matrix.

  • Model trust ledger — records model capability assertions separately from durable observed-run, benchmark, or audit evidence. Trust is derived from independently verifiable successful artifacts; self-reported claims never raise it, and open high-severity incidents restrict the model.

The trust ledger also keeps an explicit observed-run denominator, violations, and source references, so a per-model/CLI violation rate is null rather than a misleading 0% when no denominator is retained. See historical evidence and rate limits.

Future orchestrator sampling (concept only)

An orchestrator may later use the derived trust level to choose sampling frequency: unverified or provisional model/CLI pairs receive denser audit sampling, while verified pairs may be sampled less often; any open material incident restores dense sampling. This is only a measurement concept. It must not change model selection or override the routing-policy correctness floors, and a small or missing denominator must never be treated as evidence of safety.

Install

dotnet add package TokenEconomy --version 0.2.0

Dependency-free, targets net10.0. The API is pre-1.0 and may still shift — pin a version and watch releases.

Usage

using TokenEconomy;

// The seeded catalog: known Claude 4.x/5.x and OpenAI GPT-5.x/6 models.
var breakdown = ModelPriceCatalog.Default.ComputeCost(
    KnownModels.ClaudeOpus48,
    new TokenUsage(Input: 250_000, Output: 12_000, CacheRead: 40_000),
    DateTime.UtcNow);

if (breakdown.HasPrice)
    Console.WriteLine($"{breakdown.Total} {breakdown.Currency}");   // ≈ 1.57 USD
else
    Console.WriteLine(breakdown.Status);   // UnknownModel or NoPriceForDate — never a silent $0

Total is null for an unknown or unpriced model, never 0 — a missing price is always explicit. Prices carry history, so a run at an earlier timestamp is costed with the rate that was valid then.

Typed model ids

KnownModels exposes one ModelId constant for every canonical entry in the default catalog. Pricing and efficiency APIs accept these typed values directly, while ModelId converts implicitly to string for seamless use with existing callers:

var model = KnownModels.ClaudeSonnet5;
var price = ModelPriceCatalog.Default.ResolvePrice(model, DateTime.UtcNow);
var fit = ModelEfficiencyMatrix.Default.SuitabilityOf(model, TaskClass.Feature);

string providerModelId = model;
var customModel = ModelId.Of("provider-model-id");

There is deliberately no implicit conversion from string to ModelId. Dynamic or custom ids must be explicit with ModelId.Of(configurationValue). The existing string overloads remain unchanged for external input and preserve case-insensitive, dot/dash-insensitive, and alias lookup.

The constants are generated from the catalog and checked in. Maintainers can regenerate them with:

dotnet run --project tools/KnownModelsGenerator -- src/TokenEconomy/KnownModels.g.cs

The test suite renders the file in memory and compares its bytes with the checked-in output, so a catalog change without regeneration fails CI.

Benchmark price-performance

var assumption = new BenchmarkTokenAssumption(
    InputTokensPerTask: 100_000,
    OutputTokensPerTask: 10_000);
var reference = new BenchmarkCellKey(KnownModels.Gpt56Sol, EffortLevel.High);

var matrix = ModelBenchmarkMatrix.Default.Build(
    "artificial-analysis-intelligence-index-v4.3",
    assumption,
    reference,
    new DateTime(2026, 9, 11, 0, 0, 0, DateTimeKind.Utc));

var candidates = ModelBenchmarkMatrix.Default.FindCandidates(
    reference,
    "artificial-analysis-intelligence-index-v4.3",
    assumption,
    matrix.AsOfUtc);

When the publisher supplies cost per task, that cost is used. Otherwise the declared token mix is costed through ModelPriceCatalog; benchmark evidence never duplicates token rates. Evidence more than 90 days old is retained and flagged stale.

Selecting the correctness route

using TokenEconomy;

var decision = ModelRoutingPolicy.Default.RecommendCore(new()
{
    Scorecard = new()
    {
        CorrectnessRisk = 24,
        ExpectedScope = 14,
        ContextDemand = 14,
        TaskTypeAndUncertainty = 6,
        EmpiricalConfidence = 6,
        QuotaAndCostHeadroom = 0,
    },
    CorrectnessTriggers = ["persistentStateMigration"],
});

Console.WriteLine($"{decision.Route.ModelId} @ {decision.Route.ThinkingLevel}");
// gpt-6-sol @ medium — the score and migration floor both require Sol/medium.

The evaluator accepts no price catalog, cost class, or quota snapshot. Quota and cost headroom contribute only the policy's declared 0-5 intake points; hard floors are applied afterward. Provider availability may select an explicitly equivalent fallback, wait, or require an operator override, but never silently spend correctness margin.

The machine source is src/TokenEconomy/catalog/model-routing-policy.json and is validated against its JSON schema, every price-catalog model, the media catalog's CLI scopes, trust-evidence unknowns, the authoritative Markdown hash, and the deterministic generated public view. Unknown models or levels, unsupported combinations, restrictions, deprecations, and provisional evidence are returned explicitly by ModelRoutingKnowledgeBase.Resolve.

As of 2026-09-25, core Luna/medium uses gpt-6-luna, and Sol/medium and Sol/xhigh use gpt-6-sol. Terra/medium retains gpt-5.6-terra to preserve score bands and reissue steps, although it now costs more per output token than GPT-6 Sol. The operator chose the new-card GPT-6 baseline using dated prices and vendor claims; local completion cohorts remain provisional. Score weights, hard floors, and reissue rules are unchanged. Read the policy and before/after prices and the public policy view.

Every task-class view lists GPT-6 candidates and explicit GPT-5.6 fallbacks, with dated prices on the website. Historical study outcomes remain separately attributed; no GPT-6 completion rate or cost per successful card is invented. Mini/high remains the bounded-pipeline exception. Existing aliases and pins are preserved, and both GPT-6 migrations remain proposal-only: an ultra pin would require an explicit downgrade to xhigh, and local no-regression qualification is incomplete.

Migrating versioned model families

Agent Studio can consume the versioned model-migrations.v1.json catalog directly by repository path or raw URL. Its adjacent JSON Schema fixes the v1 contract, including generation order, derived cost-class movement, thinking-ladder compatibility, context change, repository evidence, and the safeAuto decision. The catalog also publishes the post-migration model sets and quota-aware alternative for chore, feature, bug, dossier, and mechanical cards.

The default strategy is latestInFamily, but automatic application is gated: same family, newer generation, same or lower known cost class, compatible ladder, and comparable no-regression evidence are all required. The consuming orchestrator must also resolve the target's dated price and current CLI availability. It applies the migration to the attempt-local route without rewriting card configuration or overriding an explicit operator pin. See the migration rules.

Status

The pricing catalog + cost API were extracted from CodingAgentRunner.Pricing (v0.5.0) into this standalone package (0.1.0); 0.2.0 adds the token-efficiency matrix + SuggestModel. TokenEconomy 0.2.0 is published on nuget.org; release operations and the one-time setup fallback are documented in docs/PUBLISHING.md, with the verified TE-1 operator handoff retained in results/TE-1-nuget-first-publish.md.

Repository layout

Path What it holds
src/TokenEconomy/ The published library. catalog/ holds the embedded price, routing, migration, task-class, and media-capability JSON.
src/TokenEconomy.Benchmarks/ CLI that executes the A/B and document-to-text benchmark runs.
tests/TokenEconomy.Tests/ xUnit suite; also the guard that the website data cannot drift from the library.
benchmarks/ Benchmark setups, fixtures and corpora, plus append-only raw results under benchmarks/results/.
docs/ Concepts, benchmark guide, publishing and repository metadata.
contexts/ Short, reusable policy blocks for agent and task-card prompts.
results/ Retained operator and backtest records.
scripts/ Release, pack, and website-data generation.
tools/ The complexity-backtest report generator.
website/ The repository's own static site.

website/ is a real part of this repository, not a mirror. It is the public documentation surface at https://agent-orchestrator.dev/token-economy/ — plain static HTML with a checked data step: scripts/generate-website-data.py regenerates website/data/*.json from the committed evidence, and CI rejects stale data, so the published charts and benchmark tables cannot drift from the library. Preview it locally on http://localhost:4340 with:

python -m http.server 4340 --directory website

See website/README.md for the content rules and website/DEPLOY.md for deployment.

Contributing

See CONTRIBUTING.md for build and test steps and the project's invariants. Issues and pull requests are welcome. Report security issues through the private process in SECURITY.md.

Agent Orchestrator ecosystem

TokenEconomy is the token-economics layer of the Agent Orchestrator family. It answers what a run costs and which model to spend the next tokens on; Agent Studio is the orchestrator that turns tasks into agent runs and is the source of the run metrics imported here; the Runner is the process and protocol layer that actually launches the coding-agent CLIs and reports the token usage this library prices; and Agent Chat is the conversation UI for those runs. See the other projects on the Agent Orchestrator website and in the agent-orc GitHub organization.

License

Apache-2.0 © 2026 Robert Mischke. See NOTICE.

About

Token Economy - token economics for LLM coding agents: model pricing with history, cost computation, and token-efficiency model selection (.NET)

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages