Skip to content

Latest commit

 

History

History
143 lines (97 loc) · 4.4 KB

File metadata and controls

143 lines (97 loc) · 4.4 KB

CodeWell V1 Evaluation Note

Last updated: 2026-05-21

Scope

This note summarizes the current V1-level internal evaluation read for CodeWell.

It is not a public benchmark claim. It is a product-facing evidence summary used to explain:

  • where CodeWell already helps
  • where it does not help
  • which task shapes require more careful interpretation

The current read is based on:

  • bundled fixture and self-evaluation coverage
  • maintained real-project evaluation coverage
  • internal Claude Code A/B runs on TS/JS wiring tasks
  • completed repeated-run analysis for strong_multi_goal prompts

V1 Positioning

The current V1 positioning is:

CodeWell is a selective navigation aid for agentic coding, not a universal acceleration layer.

What V1 already supports well:

  • local-first indexing and retrieval
  • lightweight lexical plus graph-aware context assembly
  • fast repository path-finding on navigation-heavy tasks
  • revision memory for previously failed and fixed snippet adaptations

What V1 does not claim:

  • universal speedups across all coding tasks
  • consistent gains on obvious single-file tasks
  • solved behavior on all prompt shapes with multiple plausible local repair targets

Strongest Positive Evidence

The strongest current product evidence is the navigation-heavy TS/JS wiring bucket.

Current internal read on the maintained 9-pair wiring batch:

  • navigation-heavy: elapsed ratio 0.298
  • navigation-heavy: tool-call ratio 0.606
  • navigation-heavy: token ratio 0.453

Interpretation:

  • CodeWell helps most when the agent must find the right file chain across routes, controllers, middleware, guards, validators, and support files
  • the main gain is better path-finding and workset control, not raw patch generation

Clear Boundary

The clearest negative boundary is the obvious-near-obvious bucket.

Current internal read:

  • elapsed ratio 1.392
  • token ratio 1.240

Interpretation:

  • when the likely edit target is already obvious, retrieval overhead can dominate
  • V1 should not be described as a universal speedup layer

Strong-Multi-Goal Read

The current V1 evaluation also includes a full repeated-run protocol for the strong_multi_goal task slice.

Protocol:

  • 3 A/B pairs per task
  • same prompt text across repeats
  • median used as the primary comparison
  • range and first-relevant-file stability used as diagnostics

Completed task set:

  1. adonis_kernel_auth_middleware_wiring
  2. adonis_profile_route_auth_middleware_wiring
  3. nestjs_users_admin_guard_service_wiring

Current read:

  • adonis_kernel_auth_middleware_wiring
    • first relevant file is fully stable at start/kernel.ts
    • elapsed median ratio 0.6725
    • token median ratio 0.8592
    • read: positive sample with stable first-file discovery but unstable local repair-target choice
  • adonis_profile_route_auth_middleware_wiring
    • first relevant file is fully stable at start/routes.ts
    • elapsed median ratio 1.1381
    • token median ratio 1.426
    • read: negative sample after the full three-sample protocol
  • nestjs_users_admin_guard_service_wiring
    • first relevant file is not fully stable
    • baseline mode file: src/roles/roles.guard.ts with fraction 0.6667
    • codewell mode file: src/users/users.controller.ts with fraction 0.6667
    • elapsed median ratio 0.9492
    • token median ratio 1.0013
    • read: near-neutral sample with instability already visible at the first local entry point

The main V1 lesson from this slice is:

  • some tasks are unstable after the agent reaches the right area
  • some tasks are unstable even at the first local entry point
  • therefore strong_multi_goal prompts should not be read from single runs

V1 Claim Envelope

Reasonable V1 claim:

CodeWell improves agent navigation on repository tasks with real path ambiguity, especially in route/controller/middleware/guard-style flows, while remaining intentionally lightweight and local-first.

Reasonable V1 non-claim:

CodeWell does not yet provide uniform gains on obvious tasks or on every prompt shape with multiple plausible local repair targets.

Release Use

This note is the intended summary source for:

  • README positioning
  • release-readiness discussion
  • V1 freeze notes
  • future V2 comparison baselines

Related Files

  • README.md
  • docs/AGENT_EVALUATION.md
  • docs/AGENT_EVAL_SUMMARY.md
  • docs/AGENT_EVAL_CONCLUSION.md
  • .agent-eval-wiring-temp/results/agent_eval_repeats_strong_multi_goal.json