Token economics for LLM coding agents — one tested source for model pricing (with history), run-cost computation, and token-efficiency model selection. .NET, dependency-free core.
dotnet add package TokenEconomyPackage page: TokenEconomy · GitHub releases
An agent run bills for input, output, and cache tokens at a rate that changes
over time, per model. Getting that wrong is not a rounding error: a hard-coded
price silently costs the wrong amount for every historic run, and a missing
price that defaults to 0 reports a budget as healthy while it drains.
All 25 catalog models have dated Standard API price histories and primary-source
provenance, verified through 2026-09-25. See
price history and scope for
historical corrections, cache semantics, and API cost versus subscription
consumption.
Direct code-review capability has its own benchmark category: 24 sourced measurements across five studies, with precision, known-issue coverage and original review configurations. See the review research and C# guide.
TokenEconomy keeps the prices — with their validity dates — and answers two questions from that one source: what did this run cost? and which model buys the most for the tokens I have left? Unknown is always returned as unknown.
API guide: https://agent-orchestrator.dev/token-economy/api/ — inputs, return types, reasoning effort, scores, and executable C# examples.
Docs & website: https://agent-orchestrator.dev/token-economy/ — a static
site built from website/ and deployed by
deploy-website.yml (see
website/DEPLOY.md).
Research plan: Forecast each task as a percentage of a five-hour cap defines the proposed measurement, uncertainty, repository boundaries, and GO-blocked delivery slices. Its visual explainers are part of the static site. No forecast implementation is authorised yet.
Delegation economy: the pattern guide
explains how orchestrators assign bounded work to cheaper model tiers and when
to escalate. Prompts and task cards can include the
standardized context block verbatim.
Task-cutting guide:
Dynamic workflows as a task-cutting strategy
compares one Claude workflow-sized card with small dependsOn cards using
Agent Studio token, review, retry, and gate evidence. It includes the current
Claude-only workflow constraint, Codex integration gap, TE-8 routing hook, and
a reusable AI-pattern candidate.
Prompt-enrichment analysis:
Prompt enrichment before an agent run
quantifies rule, embedding, classifier, and hybrid preprocessing against the
Agent Studio retry baseline. It specifies the auditable
enrichment-report.json contract, a Task Server integration boundary, a
selection rubric, and a standalone ai-patterns handoff.
-
Pricing catalog with history — per-model price entries keyed by
ValidFrom; a run's cost is computed with the prices valid at the run's timestamp. Historic entries are kept, never overwritten. Every listing also carries a canonicalDisplayName, so a consumer never has to maintain its own partial, drifting model-label table. -
Cost API —
ComputeCost(model, usage, atUtc)→ deterministic breakdown + total; unknown models return an explicit unknown, never a silent zero. -
Versioned model-routing policy — one embedded, schema-backed knowledge base resolves model aliases, providers/CLIs, reasoning levels, score tiers, correctness floors, workflow exceptions, restrictions, deprecations, reissues, and evidence status. The core ladder follows Agent Studio; the local policy revision adds explicit Astra, Fable 5.1, Claude Opus 5.5, GPT-6 Sol, and GPT-6 Luna support. The versioned policy is authoritative; pricing and quota cannot lower its correctness floors. See the generated knowledge view and authoritative policy.
-
Token-efficiency compatibility matrix —
SuggestModel(taskClass, budgetPressure, availableClis, atUtc)ranks only policy-selectable core models. UseModelRoutingPolicyfor correctness decisions andModelRoutingKnowledgeBase.FallbacksFor(routeId)for evidence-scoped provider fallbacks. Cost class is derived from the pricing catalog, never restated. -
Agent Studio ingest — reads each card's durable
task.jsondirectly and upserts model-run metrics by task key + run. It maps model/provider, usage, price-at-run-time, final-lane outcome signal, timestamps, project, task type, CLI and thinking level;ModelRunViewsprovides daily per-model, CLI, and project consumption/outcome views with explicit unresolved-cost counts. The filesystem contract is intentionally used over the task-server API so reporting jobs do not require a running server. -
Provider availability snapshot — captures provider/CLI probe state, independently identified quota windows, observed usage/headroom and reset, freshness, conservative warning state, and model-price coverage at one decision instant. Imported runs supply only the separately labelled rate and exhaustion projection; they never replace observed quota. Unknown, stale, suspicious, unavailable, or unpriced inputs cannot render as healthy or as a zero cost. The snapshot is routing evidence, not model selection. See the data contract and rendered snapshot.
-
Deterministic routing composition —
ModelRouter.Routecombines an upfront estimate, the versioned policy and knowledge base, benchmark and trust qualification, workflow capabilities, available CLIs, run-scoped quota/budget state, and an optional operator pin. It returns a selected model/thinking level or an explicit wait/override decision without allowing quota or cost to cross a correctness floor. See the routing API contract. -
Agent Studio admission loop —
AgentStudioTaskAdmission.PrepareAttemptstores intake features, consumes the newest classified prior outcome, routes before each attempt against a run-scoped quota snapshot, records the complete decision, and returns either an attempt-local launch route or a closed wait/override disposition. Card configuration is retained separately and is never silently rewritten. See the host integration contract. -
Completed-card economics: compare model/reasoning pairs using local retries, dated token costs, review evidence, organisation refusals, and subscription quota share through the decision API and CLI.
-
Controlled A/B benchmarks — versioned repository definitions execute the same task against model/effort variants in isolated workspaces, retain raw append-only measurements, and derive deterministic comparison reports. Start with the end-to-end benchmark methodology; the protocol background records the established-suite influences.
-
External benchmark evidence and price-performance — schema-backed, append-only benchmark definitions and model/effort measurements share one typed catalog with controlled setups.
ModelBenchmarkMatrixreturns raw or weighted score cells, dated token prices, published task metrics, declared- assumption cost, score per dollar, reference deltas, evidence age, and ranked cheaper-and-at-least-as-good candidates. It recommends; it does not route or weaken policy floors. See the public matrix and refresh recipe. -
Human-friendly language capability — dated German/English evidence for six language dimensions, typed overall and per-dimension routing constraints, and sample-cost ranking with stable ties. A validated Voice Lint importer retains source hashes and history; the capability guide and website matrix distinguish public research from measured model scores.
-
Study-backed task-class advice —
TaskClassRecommendationCatalog.Default.Recommend(taskClass)returns the ranked equivalent model/thinking set, rationale version, evidence, measured cost per outcome, and downgrade boundary used at card creation; a separateSelect(set, quotaState, atUtc)chooses concretely only from fresh quota evidence.OutcomeEfficiency.ComputeObservedcalculates tokens per accepted outcome and retry-adjusted dated list-price cost. The published task-class studies cover the complete taxonomy and the first controlled HTML/UI and source-review pilots; the concrete card still goes throughModelRouterso correctness floors win. -
Upfront task complexity — card/repository signals plus measured historical neighbours produce a versioned routing score, confidence, token/reissue forecast, and audit evidence. A host-supplied mini-model rubric is optional. See the design and backtest contract.
-
Document-to-text benchmark — a curated PDF/Word/spreadsheet/presentation hard-case corpus runs across every catalog model and derives evidence-linked, per-document-type capability records without turning failures into unsupported claims. See the capability-corpus methodology.
-
Native media capability catalog — evidence-dated image, video, music, speech, and dictation rows for Codex, Antigravity, and Claude Code are pulled from the same embedded catalog convention as pricing. Includes the retained N=4 Codex image benchmark and explicit unknown/unverified cost factors. See the capability matrix.
-
Model trust ledger — records model capability assertions separately from durable observed-run, benchmark, or audit evidence. Trust is derived from independently verifiable successful artifacts; self-reported claims never raise it, and open high-severity incidents restrict the model.
The trust ledger also keeps an explicit observed-run denominator, violations,
and source references, so a per-model/CLI violation rate is null rather than
a misleading 0% when no denominator is retained. See historical evidence
and rate limits.
An orchestrator may later use the derived trust level to choose sampling frequency: unverified or provisional model/CLI pairs receive denser audit sampling, while verified pairs may be sampled less often; any open material incident restores dense sampling. This is only a measurement concept. It must not change model selection or override the routing-policy correctness floors, and a small or missing denominator must never be treated as evidence of safety.
dotnet add package TokenEconomy --version 0.2.0Dependency-free, targets net10.0. The API is pre-1.0 and may still shift —
pin a version and watch releases.
using TokenEconomy;
// The seeded catalog: known Claude 4.x/5.x and OpenAI GPT-5.x/6 models.
var breakdown = ModelPriceCatalog.Default.ComputeCost(
KnownModels.ClaudeOpus48,
new TokenUsage(Input: 250_000, Output: 12_000, CacheRead: 40_000),
DateTime.UtcNow);
if (breakdown.HasPrice)
Console.WriteLine($"{breakdown.Total} {breakdown.Currency}"); // ≈ 1.57 USD
else
Console.WriteLine(breakdown.Status); // UnknownModel or NoPriceForDate — never a silent $0Total is null for an unknown or unpriced model, never 0 — a missing price
is always explicit. Prices carry history, so a run at an earlier timestamp is
costed with the rate that was valid then.
KnownModels exposes one ModelId constant for every canonical entry in the
default catalog. Pricing and efficiency APIs accept these typed values directly,
while ModelId converts implicitly to string for seamless use with existing
callers:
var model = KnownModels.ClaudeSonnet5;
var price = ModelPriceCatalog.Default.ResolvePrice(model, DateTime.UtcNow);
var fit = ModelEfficiencyMatrix.Default.SuitabilityOf(model, TaskClass.Feature);
string providerModelId = model;
var customModel = ModelId.Of("provider-model-id");There is deliberately no implicit conversion from string to ModelId.
Dynamic or custom ids must be explicit with ModelId.Of(configurationValue).
The existing string overloads remain unchanged for external input and preserve
case-insensitive, dot/dash-insensitive, and alias lookup.
The constants are generated from the catalog and checked in. Maintainers can regenerate them with:
dotnet run --project tools/KnownModelsGenerator -- src/TokenEconomy/KnownModels.g.csThe test suite renders the file in memory and compares its bytes with the checked-in output, so a catalog change without regeneration fails CI.
var assumption = new BenchmarkTokenAssumption(
InputTokensPerTask: 100_000,
OutputTokensPerTask: 10_000);
var reference = new BenchmarkCellKey(KnownModels.Gpt56Sol, EffortLevel.High);
var matrix = ModelBenchmarkMatrix.Default.Build(
"artificial-analysis-intelligence-index-v4.3",
assumption,
reference,
new DateTime(2026, 9, 11, 0, 0, 0, DateTimeKind.Utc));
var candidates = ModelBenchmarkMatrix.Default.FindCandidates(
reference,
"artificial-analysis-intelligence-index-v4.3",
assumption,
matrix.AsOfUtc);When the publisher supplies cost per task, that cost is used. Otherwise the
declared token mix is costed through ModelPriceCatalog; benchmark evidence
never duplicates token rates. Evidence more than 90 days old is retained and
flagged stale.
using TokenEconomy;
var decision = ModelRoutingPolicy.Default.RecommendCore(new()
{
Scorecard = new()
{
CorrectnessRisk = 24,
ExpectedScope = 14,
ContextDemand = 14,
TaskTypeAndUncertainty = 6,
EmpiricalConfidence = 6,
QuotaAndCostHeadroom = 0,
},
CorrectnessTriggers = ["persistentStateMigration"],
});
Console.WriteLine($"{decision.Route.ModelId} @ {decision.Route.ThinkingLevel}");
// gpt-6-sol @ medium — the score and migration floor both require Sol/medium.The evaluator accepts no price catalog, cost class, or quota snapshot. Quota and cost headroom contribute only the policy's declared 0-5 intake points; hard floors are applied afterward. Provider availability may select an explicitly equivalent fallback, wait, or require an operator override, but never silently spend correctness margin.
The machine source is
src/TokenEconomy/catalog/model-routing-policy.json
and is validated against its JSON schema, every price-catalog model, the media
catalog's CLI scopes, trust-evidence unknowns, the authoritative Markdown hash,
and the deterministic generated public view. Unknown models or levels,
unsupported combinations, restrictions, deprecations, and provisional evidence
are returned explicitly by ModelRoutingKnowledgeBase.Resolve.
As of 2026-09-25, core Luna/medium uses gpt-6-luna, and Sol/medium
and Sol/xhigh use gpt-6-sol. Terra/medium retains gpt-5.6-terra to preserve
score bands and reissue steps, although it now costs more per output token
than GPT-6 Sol. The operator chose the new-card GPT-6 baseline using dated
prices and vendor claims; local completion cohorts remain provisional.
Score weights, hard floors, and reissue rules are unchanged. Read the
policy and before/after prices
and the public policy view.
Every task-class view lists GPT-6 candidates and explicit GPT-5.6 fallbacks,
with dated prices on the website. Historical study outcomes remain separately
attributed; no GPT-6 completion rate or cost per successful card is invented.
Mini/high remains the bounded-pipeline exception. Existing aliases and pins are
preserved, and both GPT-6 migrations remain proposal-only: an ultra pin would
require an explicit downgrade to xhigh, and local no-regression qualification
is incomplete.
Agent Studio can consume the versioned
model-migrations.v1.json
catalog directly by repository path or raw URL. Its adjacent JSON Schema fixes
the v1 contract, including generation order, derived cost-class movement,
thinking-ladder compatibility, context change, repository evidence, and the
safeAuto decision. The catalog also publishes the post-migration model sets
and quota-aware alternative for chore, feature, bug, dossier, and
mechanical cards.
The default strategy is latestInFamily, but automatic application is gated:
same family, newer generation, same or lower known cost class, compatible
ladder, and comparable no-regression evidence are all required. The consuming
orchestrator must also resolve the target's dated price and current CLI
availability. It applies the migration to the attempt-local route without
rewriting card configuration or overriding an explicit operator pin. See the
migration rules.
The pricing catalog + cost API were extracted from CodingAgentRunner.Pricing
(v0.5.0) into this standalone package (0.1.0); 0.2.0 adds the
token-efficiency matrix + SuggestModel. TokenEconomy 0.2.0 is published
on nuget.org; release
operations and the one-time setup fallback are documented in
docs/PUBLISHING.md, with the verified TE-1 operator
handoff retained in
results/TE-1-nuget-first-publish.md.
| Path | What it holds |
|---|---|
src/TokenEconomy/ |
The published library. catalog/ holds the embedded price, routing, migration, task-class, and media-capability JSON. |
src/TokenEconomy.Benchmarks/ |
CLI that executes the A/B and document-to-text benchmark runs. |
tests/TokenEconomy.Tests/ |
xUnit suite; also the guard that the website data cannot drift from the library. |
benchmarks/ |
Benchmark setups, fixtures and corpora, plus append-only raw results under benchmarks/results/. |
docs/ |
Concepts, benchmark guide, publishing and repository metadata. |
contexts/ |
Short, reusable policy blocks for agent and task-card prompts. |
results/ |
Retained operator and backtest records. |
scripts/ |
Release, pack, and website-data generation. |
tools/ |
The complexity-backtest report generator. |
website/ |
The repository's own static site. |
website/ is a real part of this repository, not a mirror. It is the
public documentation surface at
https://agent-orchestrator.dev/token-economy/ — plain static HTML with a
checked data step: scripts/generate-website-data.py regenerates
website/data/*.json from the committed evidence, and CI rejects stale data,
so the published charts and benchmark tables cannot drift from the library.
Preview it locally on http://localhost:4340 with:
python -m http.server 4340 --directory websiteSee website/README.md for the content rules and
website/DEPLOY.md for deployment.
See CONTRIBUTING.md for build and test steps and the project's invariants. Issues and pull requests are welcome. Report security issues through the private process in SECURITY.md.
TokenEconomy is the token-economics layer of the Agent Orchestrator family. It answers what a run costs and which model to spend the next tokens on; Agent Studio is the orchestrator that turns tasks into agent runs and is the source of the run metrics imported here; the Runner is the process and protocol layer that actually launches the coding-agent CLIs and reports the token usage this library prices; and Agent Chat is the conversation UI for those runs. See the other projects on the Agent Orchestrator website and in the agent-orc GitHub organization.
Apache-2.0 © 2026 Robert Mischke. See NOTICE.