You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Publish reproducible PowerContext OFF/ON benchmarks across 3–4 agent harnesses and multiple model families, so users can assess whether the benefits hold for their own agent and model combination.
Here, agent harness means the agent host/runtime (such as Codex CLI, Goose, or Pi), distinct from the benchmark's execution and grading infrastructure.
Problem and proposed solution
The current benchmark page describes a paired SWE-bench Pro public v2 run using Codex and gpt-5.6-sol. This provides evidence for that configuration; broader coverage is needed to assess generality across agent hosts and models.
Initial coverage
Target at least three agent harnesses and three model families, with a fourth combination if practical:
Candidate agent harness
Candidate model family
Codex CLI
GPT / Sol
Goose
Qwen
Pi
GLM
Optional fourth host, such as Claude Code or OpenCode
An additional supported model family
These are candidate combinations, not claims of verified compatibility. Validate model access, PowerContext integration, unattended execution, and trace/usage capture before selecting the final matrix. Record exact host versions, provider/model IDs, reasoning settings, and integration versions.
Experimental protocol
Run paired PowerContext OFF versus ON within every selected host/model combination, using the same task IDs, grader, host configuration, model settings, permissions, and execution budgets. Keep the host's normal native capabilities enabled in both arms.
Pin the dataset, task revisions, PowerContext commit, and environment. Document precisely what ON enables and how its initial context is constructed.
Isolate each arm and trial. Permit only declared pre-task history/context; prevent leakage from reference answers, hidden tests, or prior evaluation attempts.
Use a shared task set where capabilities allow. Include real task-completion outcomes and a bounded cross-session continuation workload relevant to PowerContext. Preserve source-native grading semantics and disclose any incompatible or excluded tasks.
Repeat trials to characterize variability. Predeclare the repetition plan, include failures/timeouts and integration failures in reporting, and report paired deltas with uncertainty. A single run per arm should be identified as preliminary.
Interpret each OFF/ON delta within its own combination. Cross-combination absolute scores alone cannot identify the effects of the model or harness. Where feasible, add one overlapping configuration, such as the same model in two hosts.
Reuse existing evaluation/E2E infrastructure where appropriate, following Tracking Issue: expand benchmark coverage #1263. Start with a small runnable pilot; maintainers can then fund and execute the declared evaluation scope.
Public results
Publish a per-combination table containing task counts, OFF/ON success rates, paired changes, uncertainty, latency, and observed token usage/cost with explicit accounting boundaries and missing-data disclosure. Token reduction is an outcome to measure, not a prerequisite for reporting a combination.
Provide runnable commands, non-secret configuration manifests, dataset hashes, exact version identities, per-task results, sanitized traces, and analysis scripts. Clearly identify pilot/subset results and distinguish integration validation from measured task performance.
Acceptance criteria
Final matrix covers at least three agent harnesses and three model families, with exact versions and compatibility documented.
A small OFF/ON pilot is reproducible for each selected combination.
Shared tasks, grading, budgets, initial context, isolation, and repetition policy are documented before the main runs.
The declared evaluation scope is completed and published for every selected combination, including unsuccessful and neutral results.
Reports include per-combination paired outcomes, variability, latency, and available usage/cost data without implying unmeasured savings.
A contributor can reproduce a bounded run and regenerate its report from the published instructions and artifacts.
English and Chinese benchmark documentation links to the protocol, results, artifacts, and limitations.
Alternatives considered
More runs on one host/model improve confidence in that configuration but do not establish broader applicability.
A full host × model Cartesian matrix would be expensive. Start with 3–4 representative combinations and expand based on observed gaps.
Integration smoke tests establish compatibility but do not demonstrate better task outcomes.
Additional context
Related tracking issue: #1263. This issue focuses on cross-host/model evidence and does not require migrating the existing full LoCoMo or SWE-bench Pro workflows.
Feature description
Publish reproducible PowerContext OFF/ON benchmarks across 3–4 agent harnesses and multiple model families, so users can assess whether the benefits hold for their own agent and model combination.
Here, agent harness means the agent host/runtime (such as Codex CLI, Goose, or Pi), distinct from the benchmark's execution and grading infrastructure.
Problem and proposed solution
The current benchmark page describes a paired SWE-bench Pro public v2 run using Codex and
gpt-5.6-sol. This provides evidence for that configuration; broader coverage is needed to assess generality across agent hosts and models.Initial coverage
Target at least three agent harnesses and three model families, with a fourth combination if practical:
These are candidate combinations, not claims of verified compatibility. Validate model access, PowerContext integration, unattended execution, and trace/usage capture before selecting the final matrix. Record exact host versions, provider/model IDs, reasoning settings, and integration versions.
Experimental protocol
Public results
Publish a per-combination table containing task counts, OFF/ON success rates, paired changes, uncertainty, latency, and observed token usage/cost with explicit accounting boundaries and missing-data disclosure. Token reduction is an outcome to measure, not a prerequisite for reporting a combination.
Provide runnable commands, non-secret configuration manifests, dataset hashes, exact version identities, per-task results, sanitized traces, and analysis scripts. Clearly identify pilot/subset results and distinguish integration validation from measured task performance.
Acceptance criteria
Alternatives considered
Additional context
Related tracking issue: #1263. This issue focuses on cross-host/model evidence and does not require migrating the existing full LoCoMo or SWE-bench Pro workflows.