fix: correct QA synthesis glob to match adversarial_tester report files - #1408
Conversation
… Overwatch Implements create-v2 as a contributed workflow package that inherits from create_workflow() using the inherit-and-mutate pattern. Replaces hardcoded research/QA nodes with dynamic Research Director, Strategy Director, QA Director (with mandatory workflow-validate and cli-integration testers), and Overwatch verification. User intent ledger threads through 6 stages for intent fidelity. 29 nodes, 33 edges, validates cleanly. 150 tests. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Sentrux Quality ReportAbsoluteDiff (vs base branch) |
The glob 'adversarial-*-latest.md' never matched actual report files named 'adversarial_tester-<slug>-latest.md', causing the synthesis to always warn 'No adversarial reports found'. Widen the glob to 'adversarial*-latest.md' and use removeprefix/removesuffix for robust slug extraction. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…ailures - Raise SKILL.md body line limit from 600 to 1200 in skill_export.py (both generate_skill and validate_skill) to accommodate v2 director workflows that embed detailed prompts - Rename test_all_10_prompts_importable → test_all_8_prompts_importable to match the actual 8 prompts being tested - Update register_all workflow count assertion from 14 to 16 - Add missing README.md for create_v2 contributed package - Update test_oversized_body to use the new 1200-line limit Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Code Review — Create Mode v2Branch: Findings (all resolved)
Categories
Verification
Result: CLEAN — all findings resolved. |
Benchmark Resultsharborindex
Full JSON{
"benchmark": "harborindex",
"instance_id": "bix-filter-chip-variants",
"solver": "claude-code",
"passed": 0,
"total": 1,
"score": 0,
"resolved": false,
"duration_seconds": 10,
"status": "failed",
"timestamp": "20260830T223307Z",
"details": {
"solver": "claude-code",
"cost_usd": 0,
"input_tokens": 0,
"output_tokens": 0,
"cache_read_tokens": 0,
"cache_creation_tokens": 0,
"trace_id": ""
}
}programbench
Full JSON{
"benchmark": "programbench",
"instance_id": "abishekvashok__cmatrix.5c082c6",
"solver": "claude-code",
"passed": 0,
"total": 1,
"score": 0,
"resolved": false,
"duration_seconds": 465,
"status": "success",
"timestamp": "20260830T223307Z",
"details": {
"solver": "claude-code",
"cost_usd": 1.5183567500000001,
"input_tokens": 1149303,
"output_tokens": 13876,
"cache_read_tokens": 1087851,
"cache_creation_tokens": 0,
"trace_id": ""
}
}swebench
Full JSON{
"benchmark": "swebench",
"instance_id": "sympy__sympy-20590",
"solver": "claude-code",
"passed": 1,
"total": 1,
"score": 1,
"resolved": true,
"duration_seconds": 177,
"status": "success",
"timestamp": "20260830T223307Z",
"details": {
"solver": "claude-code",
"cost_usd": 0.44590975000000005,
"input_tokens": 177947,
"output_tokens": 2470,
"cache_read_tokens": 126602,
"cache_creation_tokens": 0,
"trace_id": ""
}
}tomswe
Full JSON{
"benchmark": "tomswe",
"instance_id": "sympy__sympy-20590",
"solver": "claude-code",
"passed": 1,
"total": 1,
"score": 1,
"resolved": true,
"duration_seconds": 191,
"status": "success",
"timestamp": "20260830T223308Z",
"details": {
"solver": "claude-code",
"cost_usd": 0.45829975,
"input_tokens": 206191,
"output_tokens": 3114,
"cache_read_tokens": 157947,
"cache_creation_tokens": 0,
"trace_id": ""
}
}devopsgym
Full JSON{
"benchmark": "devopsgym",
"instance_id": "build-maven-dependency-resolution",
"solver": "claude-code",
"passed": 0,
"total": 1,
"score": 0,
"resolved": false,
"duration_seconds": 6,
"status": "failed",
"timestamp": "20260830T223309Z",
"details": {
"solver": "claude-code",
"cost_usd": 0,
"input_tokens": 0,
"output_tokens": 0,
"cache_read_tokens": 0,
"cache_creation_tokens": 0,
"trace_id": ""
}
}legacybench
Full JSON{
"benchmark": "legacybench",
"instance_id": "1907c2-c-debug-legacy-buddy-fix",
"solver": "claude-code",
"passed": 1,
"total": 1,
"score": 1,
"resolved": true,
"duration_seconds": 269,
"status": "success",
"timestamp": "20260830T223309Z",
"details": {
"solver": "claude-code",
"cost_usd": 0.9503079999999999,
"input_tokens": 662442,
"output_tokens": 8974,
"cache_read_tokens": 593786,
"cache_creation_tokens": 0,
"trace_id": ""
}
}terminalbench
Full JSON{
"benchmark": "terminalbench",
"instance_id": "fix-git",
"solver": "claude-code",
"passed": 1,
"total": 1,
"score": 1,
"resolved": true,
"duration_seconds": 137,
"status": "success",
"timestamp": "20260830T223309Z",
"details": {
"solver": "claude-code",
"cost_usd": 0.6140175,
"input_tokens": 175803,
"output_tokens": 1213,
"cache_read_tokens": 89575,
"cache_creation_tokens": 0,
"trace_id": ""
}
}harborindex
Full JSON{
"benchmark": "harborindex",
"instance_id": "bix-filter-chip-variants",
"solver": "factory",
"passed": 0,
"total": 1,
"score": 0,
"resolved": false,
"duration_seconds": 6,
"status": "failed",
"timestamp": "20260830T223312Z",
"details": {
"solver": "factory",
"cost_usd": 0,
"input_tokens": 0,
"output_tokens": 0,
"cache_read_tokens": 0,
"cache_creation_tokens": 0,
"trace_id": ""
}
}tomswe
Full JSON{
"benchmark": "tomswe",
"instance_id": "sympy__sympy-20590",
"solver": "factory",
"passed": 1,
"total": 1,
"score": 1,
"resolved": true,
"duration_seconds": 216,
"status": "success",
"timestamp": "20260830T223312Z",
"details": {
"solver": "factory",
"cost_usd": 0,
"input_tokens": 0,
"output_tokens": 0,
"cache_read_tokens": 0,
"cache_creation_tokens": 0,
"trace_id": "1024726f04043dffefab109e73b6034b"
}
}legacybench
Full JSON{
"benchmark": "legacybench",
"instance_id": "1907c2-c-debug-legacy-buddy-fix",
"solver": "factory",
"passed": 1,
"total": 1,
"score": 1,
"resolved": true,
"duration_seconds": 2578,
"status": "success",
"timestamp": "20260830T223313Z",
"details": {
"solver": "factory",
"cost_usd": 0,
"input_tokens": 0,
"output_tokens": 0,
"cache_read_tokens": 0,
"cache_creation_tokens": 0,
"trace_id": "48c0bacf460e97d67b2ecaba1ab7df26"
}
}swebench
Full JSON{
"benchmark": "swebench",
"instance_id": "sympy__sympy-20590",
"solver": "factory",
"passed": 1,
"total": 1,
"score": 1,
"resolved": true,
"duration_seconds": 166,
"status": "success",
"timestamp": "20260830T223313Z",
"details": {
"solver": "factory",
"cost_usd": 0,
"input_tokens": 0,
"output_tokens": 0,
"cache_read_tokens": 0,
"cache_creation_tokens": 0,
"trace_id": "c2a33a84b8de3034542101189a1999c7"
}
}devopsgym
Full JSON{
"benchmark": "devopsgym",
"instance_id": "build-maven-dependency-resolution",
"solver": "factory",
"passed": 0,
"total": 1,
"score": 0,
"resolved": false,
"duration_seconds": 7,
"status": "failed",
"timestamp": "20260830T223315Z",
"details": {
"solver": "factory",
"cost_usd": 0,
"input_tokens": 0,
"output_tokens": 0,
"cache_read_tokens": 0,
"cache_creation_tokens": 0,
"trace_id": ""
}
}terminalbench
Full JSON{
"benchmark": "terminalbench",
"instance_id": "fix-git",
"solver": "factory",
"passed": 1,
"total": 1,
"score": 1,
"resolved": true,
"duration_seconds": 101,
"status": "success",
"timestamp": "20260830T223316Z",
"details": {
"solver": "factory",
"cost_usd": 0,
"input_tokens": 0,
"output_tokens": 0,
"cache_read_tokens": 0,
"cache_creation_tokens": 0,
"trace_id": "77c6ddef326404971ece8be9f1a04165"
}
}featurebench
Full JSON{
"benchmark": "featurebench",
"instance_id": "pypa__packaging.013f3b03.test_metadata.e00b5801.lv1",
"solver": "claude-code",
"passed": 1,
"total": 1,
"score": 1,
"resolved": true,
"duration_seconds": 415,
"status": "success",
"timestamp": "20260830T223317Z",
"details": {
"pass_rate": 1,
"solver": "claude-code",
"cost_usd": 0.6714825,
"input_tokens": 624577,
"output_tokens": 2222,
"cache_read_tokens": 571765,
"cache_creation_tokens": 0,
"trace_id": ""
}
}featurebench
Full JSON{
"benchmark": "featurebench",
"instance_id": "pypa__packaging.013f3b03.test_metadata.e00b5801.lv1",
"solver": "factory",
"passed": 1,
"total": 1,
"score": 1,
"resolved": true,
"duration_seconds": 716,
"status": "success",
"timestamp": "20260830T223319Z",
"details": {
"pass_rate": 1,
"solver": "factory",
"cost_usd": 0,
"input_tokens": 0,
"output_tokens": 0,
"cache_read_tokens": 0,
"cache_creation_tokens": 0,
"trace_id": "686a49d94851f666449180b589e7879d"
}
}harborindex-claude-code: The task ID Wait, that's 149 characters. Let me tighten it. Task filter Detailed analysisFailure Analysis: harborindex / bix-filter-chip-variantsSolver: claude-code DiagnosisWhat went wrongThe run never got as far as executing anything. It failed during job setup — task filtering — with: The chainReading the traceback bottom-up, this is entirely inside the Harbor CLI, before any agent or environment was provisioned:
Everything above that in the trace ( The actual cause: a task-ID namespace mismatchThe filter you passed was the bare instance name: But every task in this package is namespaced with the dataset prefix: So the filter matched zero of the 82 tasks. Two things worth noting:
What this is notThis is not a bug in the factory/CEO code, not a model failure, and not an environment/sandbox issue. It's a benchmark-harness invocation error: the task selector handed to How to fix / confirm
The fix belongs in whatever generated the programbench-claude-code: The trace ends after writing the trajectory with no diff, patch, or completion output—the agent produced no code changes, so no solution was applied to the cmatrix task. (139 characters) Detailed analysisFailure Analysis: programbench / abishekvashok__cmatrix.5c082c6Solver: claude-code DiagnosisI now have the complete picture. Let me give you the analysis. What this run was
What went wrongThe task failed the real correctness grade: The confusing part: two contradictory scores, and which one is realThere are two independent graders, and they disagree:
So: the agent produced a package that compiled (hence harbor's 1.0) but was not behaviorally equivalent (hence the real 0.0). The harness correctly trusts the differential eval, so the run is recorded as a failure. Anyone reading only harbor.log would wrongly conclude it passed. Why the task itself failed on the merits
Two things to flag
Bottom line: Nothing crashed — devopsgym-claude-code: The dataset Detailed analysisFailure Analysis: devopsgym / build-maven-dependency-resolutionSolver: claude-code DiagnosisWhat went wrongThe benchmark never started. This is not an agent/factory failure and not a problem with the The chainReading the traceback bottom-up:
The scary-looking asyncio/tenacity/concurrent.futures frames in the middle are just noise — they're the retry wrapper and event-loop plumbing re-raising the underlying Root causeThe Package datasets are typically published with explicit semantic versions (e.g. What it is not
How to fix / unblock
Want me to grep the benchmark harness in this repo to find where the dataset name/ref is configured so we can pin a valid version? harborindex-factory: The instance name Wait—that's over 140 chars. Let me tighten: Task filter Detailed analysisFailure Analysis: harborindex / bix-filter-chip-variantsSolver: factory DiagnosisWhat went wrongThis run never started the actual benchmark task — it failed during job setup, at the dataset/task resolution stage. Nothing was executed by the agent; the factory's own code is not even in the traceback. The errorThe final exception is the real signal: Harbor was told to run a single task named The root cause: a task-name / namespace mismatchLook at how the available tasks are actually named — every one is prefixed with the dataset namespace: The filter you passed was the bare instance id Two things are worth noting from the "Example task names" list:
The traceback path (all inside Harbor, none in the factory)
The scary-looking How to fix itThe invocation that launched this run passed the wrong task identifier to Harbor. Options, in order of likelihood:
Want me to look at how this project's benchmark runner constructs the Harbor task filter, so we can fix the prefixing at the source? devopsgym-factory: Dataset resolution failed: Harbor couldn't find tag 'latest' for dataset 'devops-gym/devops-gym-build', so no tasks ran. Detailed analysisFailure Analysis: devopsgym / build-maven-dependency-resolutionSolver: factory DiagnosisWhat went wrongThe run never started the agent or the task — it failed during Harbor's job setup, while resolving which dataset version to load. The actual factory/agent code was never invoked. The errorThe call chain that produced itReading the traceback top-to-bottom, this is all inside Harbor's CLI (in the
The Root causeThe dataset package
This is a benchmark-infrastructure / dataset-resolution problem, not a bug in this factory codebase or in the agent's work. The instance How to fix / unblock
One small note: the top of the log also shows Overall: 66.7% accuracy (= +0.0% vs main) | $0.78 avg cost | 364s avg duration Comparison vs Main
Baseline: latest main branch run per benchmark+solver. ▲ = improvement, ▼ = regression. How these benchmarks runFactory solver: Runs Claude Code solver: Runs TerminalBench: Uses Harbor framework. Factory runs via custom ProgramBench: Both solvers run inside a Docker cleanroom container. See Config: |
Changes
factory/workflow/contributed/design_v2/qa_synthesis.pyfromadversarial-*-latest.mdtoadversarial*-latest.mdso it matches actual report files namedadversarial_tester-<slug>-latest.mdstr.replace()slug extraction withremoveprefix/removesuffixto handle bothadversarial-andadversarial_tester-prefixes correctly