Make the buddy the tutor along the onboarding path - #208
DavidLeuter wants to merge 19 commits into
Conversation
The buddy's persona gains four clauses, each gated on the tool it depends on, so a hire with no onboarding path meets a mentor that never mentions one: - the path is the plan, and where the mentor's own suggestion disagrees with it the path wins, because a person wrote it; - a knowledge question is a tutoring moment, not an exam the mentor can shortcut: it is not told which answer is correct, and the clause says so, so "I do not have it" is the honest answer when a hire asks for it; - completing a step is asked for, never announced; - an empty AI-enhanced phase is a subject to talk through, and a step that comes out of that goes on the hire's own copy, never on the PM's blueprint. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Six things a testing session turned up, five of them the mentor's own judgement and one of them the reason it had no judgement to make. The one that matters most: "where am I?" had two plausible answers — the path a person wrote, and the work pool, metrics and ledger that say what is true right now — and which one it landed on depended on nothing the hire could see. That is the two-systems problem moved inside one conversation rather than solved, so the routing is now stated: path questions to the path, always; the work pool only when the path has nothing open or the current step is asking for real work; metrics and the ledger answer how it is going and never what comes next. Mounted only where both exist. Also: - A LOCKED item is never agreed to, however directly asked. It said yes to a step the hire's own page refuses to open. - Items are named by the number their page prints and linked, so "want to take the check?" arrives as something clickable. - Part of a step is a line of its checklist, not the step. - A refusal that came back NOT PROPOSED has no button: it had been telling hires to click something that was never rendered. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two causes, one symptom: the mentor kept telling hires to confirm something that was never on their screen. **The search budget.** Only search-only hops consume it — a hop that asks for a backend tool returns immediately — so a model that searched four times before deciding to act never got to act: it fell through to the forced final answer, which is generated with no tools, while the persona in front of it still said "offer `add_path_step`". It took that at face value. The forced turn is now told plainly that it has no tools and can offer nothing, and the budget is 6 rather than 4 so a thorough answer still leaves room to act. **The mechanics were never stated.** The persona said what to offer and never that offering *is* a tool call. It now says so, that a call missing an argument is refused and shows no button, and that describing a button is not making one. Also: an item is written as a linked number in a literal shape (`[#3](/onboarding/...) "title"`) rather than as a description of one — a small model copies a pattern far more reliably than it interprets an instruction about formatting. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
There is no skip action and there should not be one: a skip is a request the PM decides, made on the step's own page. The persona now says so, so the mentor neither promises one nor ignores the ask. Adds a test that a spent search budget sends the model the no-tools notice and keeps it out of the conversation, and counts complete_task among the tools a hire without a path never hears about. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The mentor brings up a step whose checklist is done but that was never finished, once, with complete_step; and after a wrong answer it teaches first and offers one refresher step when the gap is bigger than one explanation. The link example follows the new /onboarding?step= shape. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Replaces "skipping is not yours to offer" with a clause gated on request_skip: ask why first, help put the reason into a sentence or two a PM can decide on, send it in the hire's words, and never promise it will be accepted. A step waiting on that decision is not pushed, and finishing it is not offered without saying that it withdraws the request. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The numbers are the page's count, not an order: an item opens once what it comes after is done, and several can open at once. Added steps are placed with waits_on and unlocks rather than dropped at the end. A ready-to-close step is answered where-they-are first, never button first and never with a lecture. Narrowing a question's options down counts as hinting. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The persona still carried the old ramp: a mentor guiding a hire "from their first day to doing real work", Task 0 as an action, and a closing line that called everything else "the path to" a first contribution. Onboarding is the path a PM's blueprint prescribes; the buddy tutors along it. - Identity reworded around the onboarding, not a first contribution. - Path clauses come first; the arrival clause follows as setup, not as part of the onboarding. - Routing: "how is my onboarding going" goes to the path; the metrics answer how their work is going and never how far along they are. - claim_task_zero removed from the action tools. Refs #311 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A hire with issues sitting in the pool asked their buddy whether there was anything they could do and was told there was not. The routing clause is why: it said the suggested work came up "only when the path has nothing open", so a half-finished path read as a closed door. That is a gate nobody designed, in the one place the hire cannot see it -- and picking work up is not a reward for getting far enough along a curriculum. So the pool is its own road now, open from day one. An explicit ask for something to pick up goes straight to the suggested work, "what should I work on" is treated as the genuinely ambiguous question it is and gets both answers, and questions about the onboarding still belong to the path. Two clauses follow the hire past the claim, which is where the mentor used to go quiet: the task packet and the docs for the work itself, and -- for the path -- reading what a step points at rather than only pointing at it. A step that says "read issue 123" names something the corpus holds. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Keeps both sides: the path tutor from #311 and what dev brought in (team-mode buddy, board checklist actions, PATH_STEP card, the reworked onboarding journey). Task 0, the ramp and the path-to-first-contribution card stay retired; dev code that still reached for them was adapted. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Dev's mode tests pinned the hire persona by its old opening line, "the mentor who guides a new hire". #311 rewrote that line to the tutor along the path, so the default-mode tests failed and the team-mode negative check passed for the wrong reason. They now look for "the tutor who guides a new hire". Refs #311 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The path read now says where the hire stands on a path whose phases run side by side. The persona says it too, so the rule holds before the tool is read and when the hire names a phase the read has not seen them start: which open phase comes next is the hire's choice, never the lowest number. Refs #311 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Testing `open_orientation` from the buddy failed on a local model: the packet came back as invalid JSON, and orientation gave up on the first parse error, so the hire got no packet at all. Phase assembly already asks the model once to correct its own JSON; orientation now does the same before it skips. Refs #311 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Same failure as orientation, in two more places the buddy reaches: - Grading a short-text answer: an unreadable reply marked every answer in it "could not be graded", which the backend records as wrong. A hire with the right answer was told it was wrong. - A diagram card: one broken quote cost the whole diagram. Both now ask the model once to correct its own JSON before giving up, the way phase assembly and orientation do. Refs #311 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Grading no longer builds a correction prompt it will never send, the diagram retry uses a _correction_prompt helper like orientation and phase assembly, and the question clause tells the mentor to start from the answer the path read now shows. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
"Offer flag_to_pm as the last resort" was the only thing the persona said about escalating, and the mentor applied it to the hire too: an explicit "flag this to my PM" was refused as not being a PM matter. A new clause, mounted with the tool, says their asking is the reason. The add-step clause gets the same exception for a step they ask for themselves. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Follow-up fix (8f6adc3): offer While testing, the buddy refused a direct "flag this to my PM". The only thing the persona said about escalating was "offer Changes:
The matching tool-description change is in SprintStartProject/sprintstart-backend#261 (00203027). Verified manually. ruff, format and pyright are clean; 999 tests pass. |
| raw = llm.generate(messages) | ||
| except LLMUnavailableError as exc: |
There was a problem hiding this comment.
The retry only handles JSON/Pydantic parsing failures, but _Payload.results currently defaults to an empty list. Therefore a syntactically valid but incomplete response such as {}, or one that omits one requested answer ID, reaches this return without retrying. The caller then reports every missing grade as incorrect even though the model never graded it.
Please make results required and validate that the response contains exactly one result for every requested ID before returning. Missing, duplicate, or unexpected IDs should raise ValueError so the existing correction round is used. For example:
payload = _Payload.model_validate_json(extract_json_object(raw))
graded_by_id = {item.id: item for item in payload.results}
expected_ids = {item.id for item in to_grade}
if (
len(payload.results) != len(to_grade)
or set(graded_by_id) != expected_ids
):
raise ValueError("grading response does not match requested answer IDs")
return graded_by_idAlso change _Payload to:
class _Payload(BaseModel):
results: list[_GradedItem]
Summary
Persona and agent changes so the buddy tutors along the hire's onboarding path instead of acting as a second onboarding. Also: grading, orientation and diagrams now ask the model once more for corrected JSON instead of failing on its first broken reply.
Refs SprintStartProject/Wiki#311
Changes
Persona (
buddy_persona.py)search_docs, tutoring on knowledge questions without giving the answer away (starting from the hire's last wrong answer when the path shows it), closing steps that are READY TO CLOSE, checklist lines, adding steps with bothwaits_onandunlocks, skip requests, and routing between the path, the work pool, metrics and the ledger.open_orientation+search_docs)._NO_BUTTON_CLAUSE: the buddy offers something by calling its tool, and aNOT PROPOSEDresult means no button was shown.Agent (
buddy_agent.py)_MAX_STEPSgoes from 4 to 6.Retry broken JSON once
grading.py,orientation.pyanddiagram.pysend one correction round when the model's JSON can't be parsed. Before, grading marked the answer wrong, and orientation or diagram gave up.Testing
ruff check,ruff format --check: greenpyright src/: cleanpytest: 988 passed. 10 local failures in ingestion code this PR doesn't touch. They fail the same way on a cleanorigin/dev: 9 tree-sitter tests under my local Python 3.14 (CI runs 3.12), and a timing threshold intest_ingest_run_processes_artifacts_concurrently(elapsed < 0.5).Notes
dev.Related PRs
🤖 Generated with Claude Code