This document narrows the next execution phase for CodeWell.
It is intentionally not an agent-expansion roadmap. Agent-facing APIs should remain modular and stable, but deeper agent capability work is deferred until the local product path is more mature.
Use this section after clearing chat history.
Current phase status:
- detached-library trust/onboarding/recovery work is largely complete
- UI productization for read-only ops/health/repair/provenance is largely complete
- command suggestion logic is now shared across CLI, UI, and detached-library status surfaces
- unified intake/dropzone import now exists under the detached library root
- intake now classifies ZIP archives, source folders, bare code files, papers, and loose documents
- managed code imports now receive read-only protection, plus best-effort Windows ACL hardening
- intake can auto-index imported code workspaces immediately after protection
- intake-imported papers/documents can now surface through
contextand MCP as lightweight relevant references without requiring manuallink-doc - the main unresolved product risk is still JS/TS evaluation depth on non-trivial real projects
- the main unresolved document-product gap is still PDF-specific paper extraction quality
- early Claude Code A/B evidence now shows CodeWell helps most on navigation-heavy repository tasks, while some obvious single-file tasks show weak or negative gain
Do not lose these decisions:
- raw source trees and original ZIP archives must remain immutable inputs
- derived artifacts, manifests, ingest history, and repair audit stay under detached library storage
- the external entry point should stay unified around managed intake, not many parallel import flows
- desk-pet / drag-and-drop UI is intentionally deferred and should be treated as a future plugin
- do not prioritize visual shell work over retrieval/context quality
- do not add embeddings/reranking before stronger retrieval evidence exists
- prefer shared helpers over duplicated command-generation logic across CLI/UI
- reuse one document-reference shape across CLI context, MCP, and future UI instead of inventing a second parallel document schema
Immediate next execution order:
- improve paper-quality extraction for PDF imports: title, abstract, and keywords
- make intake-imported documents project/file/symbol-aware with stronger semi-automatic linking
- extend the maintained JS/TS real-project set with another framework-heavy backend target
- inspect graph precision/recall misses from the maintained evaluations
- add only conservative parser/graph improvements justified by those misses
These items come directly from the current Claude Code A/B runs.
The main takeaway is not that retrieval is broadly weak. The main takeaway is that retrieval needs to become more selective. On ambiguous navigation-heavy tasks, CodeWell helps. On obvious single-file tasks, retrieval overhead can outweigh the benefit.
Goal: do less work when the main file is already obvious.
Why this matters:
- negative-gain tasks so far were short bootstrap or schema tasks
- the model could identify the main file immediately
- additional retrieval and graph expansion did not reduce ambiguity enough to pay for itself
Needed behavior:
- if search confidence is already high and one obvious file dominates, reduce graph expansion
- treat short high-confidence entry files as a stop signal, not an invitation to expand neighbors
- prefer a minimal answer over a broad context pack when uncertainty is low
Likely touchpoints:
src/codewell/context.pysrc/codewell/search.py
Goal: adapt retrieval strategy to the kind of task being asked.
Why this matters:
- CodeWell performs best on middleware, redirect, kernel, and multi-file entrypoint flows
- it performs worse on obvious schema or bootstrap tasks
Needed behavior:
- bias toward graph expansion for:
- middleware
- kernel
- route registration
- redirect / guard / auth wiring
- bias toward direct top-hit retrieval with minimal expansion for:
- obvious schema tasks
- short bootstrap tasks
- local single-controller or single-file tasks
This does not need a heavyweight classifier first. A conservative keyword and path-shape heuristic would already be useful.
Likely touchpoints:
src/codewell/context.pysrc/codewell/search.pysrc/codewell/models.py
Goal: improve exactly the kinds of path-finding CodeWell is supposed to help with.
Why this matters:
- strongest gains came from navigation-heavy tasks, not local code reasoning
- this is where additional retrieval quality will compound product value
Highest-value relationships to deepen:
- route registration -> controller / handler
- kernel / bootstrap registration -> middleware / interceptor / pipe implementation
- auth config -> redirect target / protected route / login path
- route entry -> feature component -> API client
Constraint:
- keep these relationships explicit and static where possible
- avoid broad framework inference unless evaluation evidence justifies it
Goal: stop pulling in files that are technically related but not likely edit targets.
Why this matters:
- several broad runs drifted into nearby but non-target files
- even when the final answer was correct, extra neighbor exploration consumed time and tokens
Needed behavior:
- lower ranking for shared support files once the feature-local path is already strong
- lower ranking for generic helpers, shared UI, and broad framework support modules unless query terms explicitly ask for them
- prefer same-feature files over same-layer but different-feature files
Likely touchpoints:
src/codewell/context.py
Goal: help the agent decide faster where to edit and where not to edit.
Why this matters:
- some tasks still had multiple plausible patch sites even after the right path was found
- the current output is useful, but not always decisive enough
Needed additions to agent-facing output:
- likely primary edit file
- likely secondary files
- likely non-target neighbor files
- short rationale for why the primary file was selected
This should remain concise. The point is to reduce drift, not generate a long explanation.
Likely touchpoints:
src/codewell/context.pysrc/codewell/models.py
Goal: avoid optimizing the product against tasks that do not represent its intended value.
Why this matters:
- obvious single-file tasks are still valid product behavior checks
- but they are poor headline evidence for navigation-oriented retrieval value
Recommended policy:
- keep obvious schema/bootstrap tasks in evaluation
- but classify them separately from navigation-heavy tasks
- do not use them as the dominant signal when deciding whether retrieval strategy is improving
This section turns the retrieval findings into a concrete implementation program. The target is not "more retrieval." The target is selective retrieval with better payoff per token, per tool call, and per second.
Optimize for this outcome:
When the answer path is ambiguous, CodeWell should reduce search cost and drift. When the answer path is obvious, CodeWell should stay lightweight and get out of the way.
That means the system should become better at both:
- expanding aggressively in the right cases
- refusing to expand in the wrong cases
Any retrieval improvement should satisfy all of these:
- locally computable
- statically explainable
- measurable against task buckets
- cheap enough for default use
- removable or reversible if evaluation does not improve
Avoid improvements that are:
- opaque
- embedding-dependent
- expensive by default
- impossible to audit from returned output
Goal: install a lightweight decision layer between lexical hit quality and graph expansion.
Core idea:
- first score direct hits
- then estimate ambiguity
- only expand when ambiguity is above a threshold
Proposed implementation:
-
Add an internal
retrieval_modedecision inbuild_context_packdirectfocused_expandnavigation_expand
-
Compute a cheap ambiguity score from signals already available:
- score gap between top hit and second hit
- path specificity of top hit
- whether top hit is an obvious entry file such as
main.ts,schema.ts,paths.ts - number of query terms matched in file path vs body only
- whether the hit is same-feature local vs generic shared/support
-
Route behavior by mode:
direct: keep only top file(s), minimal neighbors, minimal tracesfocused_expand: allow one-hop expansion with strict capsnavigation_expand: allow richer one-hop expansion and bounded second-hop candidates
Why this is advanced:
- it adds adaptive behavior without a heavyweight model
- it treats retrieval like a decision problem rather than a fixed pipeline
Why this is efficient:
- uses signals already present in ranking and selection
- reduces work on easy tasks instead of increasing work everywhere
Suggested touchpoints:
src/codewell/context.pysrc/codewell/search.pysrc/codewell/models.py
Acceptance criteria:
- negative-gain tasks should show lower token and tool usage
- strong-gain tasks should preserve or improve recall
- debug output should make the chosen mode visible
Goal: infer the kind of navigation problem from the query before expansion.
Core idea:
Use a small explicit heuristic layer to classify queries into shapes such as:
bootstrap_wiringmiddleware_wiringredirect_guardroute_entryvalidation_schemaservice_flowsingle_file_local
Proposed approach:
-
Build a keyword and path-shape detector:
middleware,kernel,guard,interceptor,piperedirect,protected,login route,authschema,validation,dtobootstrap,main,startup
-
Combine query-shape priors with file-shape priors:
start/kernel.tsstrongly matches middleware wiringsrc/main.tsstrongly matches bootstrap wiring*.schema.tsstrongly matches validation
-
Tune expansion strategy by query shape:
- middleware and route-entry shapes get stronger graph/path expansion
- validation and bootstrap shapes get direct-hit bias
Why this is advanced:
- retrieval behavior becomes intent-sensitive instead of globally uniform
Why this is efficient:
- no model call
- cheap heuristics
- directly informed by observed gain buckets
Acceptance criteria:
- task bucket performance becomes more separated in the expected direction
- fewer unnecessary neighbors appear in bootstrap/schema tasks
Goal: improve static graph coverage exactly where CodeWell has proven value.
Do not widen the graph indiscriminately. Add only relationships with clear evaluation demand.
Priority relationship upgrades:
-
registration edges
- bootstrap -> global pipe / interceptor / middleware class
- kernel -> middleware implementation
- route mount -> router / controller / handler
-
auth/navigation edges
- protected route -> redirect target
- auth config -> login / logout / session paths
- route entry -> feature component -> API client
-
validation wiring edges
- route -> validator / schema / middleware
- DTO / schema -> request handling call sites
Implementation guidance:
- prefer explicit parser-backed extraction first
- add framework-specific heuristics only where task evidence already exists
- each new edge type must have at least one task-level regression test
Acceptance criteria:
- strong-gain tasks improve further or become more stable
- graph-size growth stays bounded
- explanation quality improves with each added edge type
Goal: reduce technically related but operationally distracting files.
Current drift patterns to suppress:
- shared UI primitives
- generic helpers
- broad framework support modules
- same-layer files outside the active feature path
Proposed scoring improvements:
-
feature-locality boost
- prefer files sharing a deeper path prefix with the top hit
-
shared-support penalty
- stronger penalty for
shared,common,utils, generic support modules unless query terms explicitly request them
- stronger penalty for
-
already-solved-path penalty
- if top hit is high-confidence and feature-local, reduce neighbor eligibility for broad support files
-
path-role penalty
- penalize paths that are usually non-edit support in the current task shape
Acceptance criteria:
- fewer non-target neighbors in context packs
- fewer same-run excursions into shared or framework support files
Goal: make retrieval output easier for agents to act on immediately.
The current agent_brief should evolve into a decision-oriented summary.
Recommended additions:
primary_edit_candidatessecondary_supporting_filesnon_target_neighborsselection_confidenceretrieval_mode- short
why_this_filerationale
Keep this concise. It should steer the agent, not narrate the whole scoring process.
Recommended output pattern:
- one likely edit target
- one or two supporting files
- one sentence saying why this is the likely patch area
Acceptance criteria:
- fewer runs drift after reaching the right feature
- shorter time from first relevant file to first edit
Implement in this order:
- Phase A: Precision Governor
- Phase D: Noise Suppression Layer
- Phase E: Agent-Facing Decision Output
- Phase B: Query Shape Heuristics
- Phase C: Navigation Edge Deepening
Reasoning:
- A and D can reduce negative-gain behavior quickly
- E improves agent usefulness without changing graph semantics
- B should be added only after the governor exists
- C has the highest upside but also the highest maintenance cost, so it should be guided by the first rounds of precision tuning
Do not measure retrieval as one aggregate number.
Track by task bucket:
navigation_heavymixedobvious_single_file
For each bucket, track:
- median elapsed seconds
- median tool calls
- median token usage
- first relevant file hit rate
- final patch in expected area
- wrong-context event rate
Primary success rule:
- navigation-heavy tasks improve materially
- obvious-single-file tasks do not regress badly
Secondary success rule:
- explanation quality and output determinism improve
Use these as low-risk first experiments:
- add
retrieval_modewith direct/focused/navigation - add stronger stop conditions for obvious top hits
- add a stronger penalty for shared support files when confidence is already high
- add
primary_edit_candidatesto the context pack output
These four changes together should produce useful signal before any deeper graph work.
Control:
- only reduce expansion when ambiguity is low
- test each change against navigation-heavy tasks first
Control:
- keep the heuristic vocabulary short
- require task evidence before adding a new task shape
- document each heuristic family and why it exists
Control:
- optimize for fewer but stronger recommendations
- put diagnostics behind a debug mode
Control:
- no new edge family without a concrete evaluation miss motivating it
- no framework-specific inference without a bounded task set to validate it
CodeWell already has the core local loop in place:
- immutable-raw indexing via detached-library mode
- unified ingest for local folders, GitHub URLs, and local ZIP archives
- unified managed intake for external files, archives, papers, and code drops
- SQLite lexical search, trace, and graph-aware context packs
- lightweight document retrieval through both manual links and intake-imported references
- revision memory for failed reuse and verified fixes
- detached-library health, repair, audit, admin overview, and local UI
- MCP access to the same local retrieval and maintenance surface
This means the next bottlenecks are no longer basic capability gaps. The main risks now are:
- weak first-run ergonomics
- uneven trust and boundary communication
- limited retrieval evaluation coverage outside Python
- insufficient operational clarity for real users
Do not prioritize these items in the current phase:
- expanding CodeWell into a full autonomous repository agent
- widening MCP surface area only for future agent orchestration
- adding embeddings, rerankers, or hosted dependencies before stronger evaluation evidence
- adding more languages before JS/TS evaluation depth catches up
These remain valid future directions, but they are not the best next investment.
These items directly affect product trust, onboarding quality, and whether the local-first model feels complete.
Goal: make detached-library mode the clearest trusted path, not just an advanced option.
Deliverables:
- add a dedicated initialization command such as
codewell init-library - create or validate a recommended library layout under one root
- clearly separate:
- user raw sources
- derived workspace data
- GitHub cache
- archive cache
- repair audit and future diagnostics
- print next-step instructions after initialization
Done when:
- a user can create a clean library root without reading architecture docs
- the CLI explains where raw sources stay and where derived artifacts go
- detached mode becomes the obvious default recommendation in docs
Suggested modules:
src/codewell/cli.pysrc/codewell/library.py- new helper module if layout bootstrap logic becomes non-trivial
Suggested tests:
- CLI init command creates expected directories
- rerunning init is idempotent
- help text explains raw/derived separation
Goal: make the immutability guarantee explicit and easy to verify.
Deliverables:
- add a dedicated README section for raw-source safety
- document detached-library lifecycle:
- what CodeWell reads
- what it writes
- where it writes
- what can be deleted and rebuilt
- explain the role of manifest, derived DB, ingest history, repair audit, and revision memory
Done when:
- a cautious user can answer “will this modify my original code?” from the main docs alone
- docs clearly distinguish durable memory from replaceable cache-like artifacts
Suggested files:
README.mddocs/PROJECT_GUIDE.mddocs/ARCHITECTURE_V2.md
Goal: make the “just throw code in” flow predictable even when it fails.
Deliverables:
- tighten CLI feedback for each source type:
- local folder
- GitHub URL
- local ZIP archive
- show normalized ingest stages and failure points more clearly
- improve recovery guidance after:
- malformed archive
- missing extracted root
- duplicate/repeated import
- failed index pass
- document when reindexing, repair, or manual review is the right next step
Done when:
- failures point to one concrete next command
- users can tell whether a problem is source acquisition, materialization, or indexing
Current status:
- failed ingest runs are stored with stage history
- CLI recovery guidance exists for local folders, GitHub URLs, and ZIP archives
- empty ZIP archives are rejected explicitly
- unsafe ZIP paths are rejected explicitly
- zero-supported-source runs now emit strong warnings
- UI now exposes inspect/suggested commands for failed ingest runs
- broader duplicate/repeated import guidance is still future work if real user reports justify it
Suggested modules:
src/codewell/ingest.pysrc/codewell/archive_ingest.pysrc/codewell/cli.pysrc/codewell/library_status.py
Suggested tests:
- more ingest failure-path CLI assertions
- recovery text for broken archive and stale manifest scenarios
Goal: stop relying on mostly fixture-level confidence for non-Python retrieval quality.
Deliverables:
- add stable JS/TS project evaluation tasks beyond parser fixtures
- cover queries involving:
- route entrypoints
- barrel chains
- re-export chains
- object methods
- class methods
- tests
- CLI/command entrypoints
- compare precision and recall before changing ranking behavior again
Done when:
- retrieval changes can be judged against task-level JS/TS evidence
- parser and graph improvements are not accepted on intuition alone
Suggested files:
evaluations/docs/EVALUATION.mdtests/test_evaluate_project.py
These are high-value quality improvements, but they can follow the must-do items.
Goal: let advanced users tune context size and bias without changing internals.
Potential additions:
- neighborhood count limit
- caller/callee definition cap
- include or skip revision memory
- bias toward tests, routes, or commands
- compact/default/debug response presets
Done when:
- context packs can be made smaller or more diagnostic intentionally
- retrieval debugging does not require code edits
Goal: make ranking and graph expansion easier to trust.
Potential additions:
- stronger “why selected” summaries at pack level
- explicit “why expanded” and “why excluded” diagnostics
- optional debug view for ranking or graph thresholds
Done when:
- maintainers can inspect a miss without reading SQL and ranking logic
- future reranking work has an auditable baseline
Goal: make long-term fix memory more reusable and less noisy.
Potential additions:
- snippet ID naming guidance
- verification-state policy guidance
- applicability note templates
- duplicate or superseded revision handling
Done when:
- revision memory can grow without turning into an unstructured patch pile
Goal: move from “useful internal page” to “real local ops console”.
Completed in this phase:
- library-root onboarding hints
- repair-plan inspection panel
- audit filters
- manifest/source provenance view
- import history and failure drill-down
- clearer next-step command presentation in workspace-health and repair views
- source-kind markers plus copy-path affordances for provenance and boundary views
- failed-ingest inspect/suggested commands in recent ingest runs
- user-facing terminology standardized around inspect command vs suggested command
Done when:
- a user can inspect library health without switching immediately to CLI
These are valid but deliberately postponed.
Keep interfaces modular, but do not optimize the whole roadmap around agent orchestration yet.
For now:
- preserve stable data models
- avoid leaking implementation details across modules
- leave extension room for compact/default/debug response modes
Only revisit after:
- stronger JS/TS evaluations exist
- lexical plus graph retrieval misses are well understood
- baseline precision and latency tradeoffs are documented
Only revisit after current Python and JS/TS coverage is better evaluated and stabilized.
Execute in this order:
- library initialization UX
- raw/derived trust documentation
- ingest feedback and recovery guidance
- JS/TS task-level evaluation expansion
- context-pack control surface
- retrieval explainability improvements
- revision-memory discipline
- UI productization
This order keeps the project aligned with its declared advantages:
- simple enough for non-experts
- local-first
- lightweight
- fast
- precise
- memory-bearing
- evolvable without mutating source inputs
Even while agent work is postponed, keep these interfaces clean:
ingest: source normalization and acquisitionlibrary: raw/derived/manifest path resolutionstatus/admin/repair: operational health and safe maintenanceretrieval: search, trace, context-pack assemblymemory: failures, revisions, applicability, verificationmcp: structured external protocolui: human inspection surface
Do not couple future UI or agent features directly to storage internals when a model-layer return type can carry the boundary instead.
This phase is complete when:
- detached-library setup feels first-class instead of expert-only
- documentation makes raw-source immutability obvious
- ingest failures are recoverable without reading code
- JS/TS retrieval quality is measured at task level, not mostly inferred from fixtures
- context packs are easier to control and easier to audit