Last updated: 2026-05-21
This note summarizes the current V1-level internal evaluation read for CodeWell.
It is not a public benchmark claim. It is a product-facing evidence summary used to explain:
- where CodeWell already helps
- where it does not help
- which task shapes require more careful interpretation
The current read is based on:
- bundled fixture and self-evaluation coverage
- maintained real-project evaluation coverage
- internal Claude Code A/B runs on TS/JS wiring tasks
- completed repeated-run analysis for
strong_multi_goalprompts
The current V1 positioning is:
CodeWell is a selective navigation aid for agentic coding, not a universal acceleration layer.
What V1 already supports well:
- local-first indexing and retrieval
- lightweight lexical plus graph-aware context assembly
- fast repository path-finding on navigation-heavy tasks
- revision memory for previously failed and fixed snippet adaptations
What V1 does not claim:
- universal speedups across all coding tasks
- consistent gains on obvious single-file tasks
- solved behavior on all prompt shapes with multiple plausible local repair targets
The strongest current product evidence is the navigation-heavy TS/JS wiring bucket.
Current internal read on the maintained 9-pair wiring batch:
navigation-heavy: elapsed ratio0.298navigation-heavy: tool-call ratio0.606navigation-heavy: token ratio0.453
Interpretation:
- CodeWell helps most when the agent must find the right file chain across routes, controllers, middleware, guards, validators, and support files
- the main gain is better path-finding and workset control, not raw patch generation
The clearest negative boundary is the obvious-near-obvious bucket.
Current internal read:
- elapsed ratio
1.392 - token ratio
1.240
Interpretation:
- when the likely edit target is already obvious, retrieval overhead can dominate
- V1 should not be described as a universal speedup layer
The current V1 evaluation also includes a full repeated-run protocol for the strong_multi_goal
task slice.
Protocol:
3A/B pairs per task- same prompt text across repeats
- median used as the primary comparison
- range and first-relevant-file stability used as diagnostics
Completed task set:
adonis_kernel_auth_middleware_wiringadonis_profile_route_auth_middleware_wiringnestjs_users_admin_guard_service_wiring
Current read:
adonis_kernel_auth_middleware_wiring- first relevant file is fully stable at
start/kernel.ts - elapsed median ratio
0.6725 - token median ratio
0.8592 - read: positive sample with stable first-file discovery but unstable local repair-target choice
- first relevant file is fully stable at
adonis_profile_route_auth_middleware_wiring- first relevant file is fully stable at
start/routes.ts - elapsed median ratio
1.1381 - token median ratio
1.426 - read: negative sample after the full three-sample protocol
- first relevant file is fully stable at
nestjs_users_admin_guard_service_wiring- first relevant file is not fully stable
- baseline mode file:
src/roles/roles.guard.tswith fraction0.6667 - codewell mode file:
src/users/users.controller.tswith fraction0.6667 - elapsed median ratio
0.9492 - token median ratio
1.0013 - read: near-neutral sample with instability already visible at the first local entry point
The main V1 lesson from this slice is:
- some tasks are unstable after the agent reaches the right area
- some tasks are unstable even at the first local entry point
- therefore
strong_multi_goalprompts should not be read from single runs
Reasonable V1 claim:
CodeWell improves agent navigation on repository tasks with real path ambiguity, especially in route/controller/middleware/guard-style flows, while remaining intentionally lightweight and local-first.
Reasonable V1 non-claim:
CodeWell does not yet provide uniform gains on obvious tasks or on every prompt shape with multiple plausible local repair targets.
This note is the intended summary source for:
- README positioning
- release-readiness discussion
- V1 freeze notes
- future V2 comparison baselines
README.mddocs/AGENT_EVALUATION.mddocs/AGENT_EVAL_SUMMARY.mddocs/AGENT_EVAL_CONCLUSION.md.agent-eval-wiring-temp/results/agent_eval_repeats_strong_multi_goal.json