Skip to content

Make the buddy the tutor along the onboarding path - #208

Open
DavidLeuter wants to merge 19 commits into
devfrom
feature/311-buddy-onboarding-tutor
Open

DavidLeuter wants to merge 19 commits into
devfrom
feature/311-buddy-onboarding-tutor

Conversation

@DavidLeuter

@DavidLeuter DavidLeuter commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Persona and agent changes so the buddy tutors along the hire's onboarding path instead of acting as a second onboarding. Also: grading, orientation and diagrams now ask the model once more for corrected JSON instead of failing on its first broken reply.

Refs SprintStartProject/Wiki#311

Changes

Persona (buddy_persona.py)

  • New identity: a tutor along the path, not "first day to real work". The old Task 0 / ramp wording is gone, and a test keeps it out.
  • New clauses, each shown only when its tool is available: reading the path, numbered and linked references, locked items, phases that run side by side (the hire picks, not the lowest number), walking through material with search_docs, tutoring on knowledge questions without giving the answer away (starting from the hire's last wrong answer when the path shows it), closing steps that are READY TO CLOSE, checklist lines, adding steps with both waits_on and unlocks, skip requests, and routing between the path, the work pool, metrics and the ledger.
  • The work pool is open from day one. The buddy stays with the hire after they claim a task (open_orientation + search_docs).
  • _NO_BUTTON_CLAUSE: the buddy offers something by calling its tool, and a NOT PROPOSED result means no button was shown.

Agent (buddy_agent.py)

  • _MAX_STEPS goes from 4 to 6.
  • The forced final answer now gets a system note that no tools are available. Before, it promised buttons that didn't exist.

Retry broken JSON once

  • grading.py, orientation.py and diagram.py send one correction round when the model's JSON can't be parsed. Before, grading marked the answer wrong, and orientation or diagram gave up.

Testing

  • ruff check, ruff format --check: green
  • pyright src/: clean
  • pytest: 988 passed. 10 local failures in ingestion code this PR doesn't touch. They fail the same way on a clean origin/dev: 9 tree-sitter tests under my local Python 3.14 (CI runs 3.12), and a timing threshold in test_ingest_run_processes_artifacts_concurrently (elapsed < 0.5).

Notes

  • The clauses are keyed to tool names, so they only show up once the backend mounts the new tools. That makes the merge order with the backend PR uncritical.
  • Up to date with dev.

Related PRs

🤖 Generated with Claude Code

DavidLeuter and others added 18 commits September 13, 2026 16:13
The buddy's persona gains four clauses, each gated on the tool it depends on, so a
hire with no onboarding path meets a mentor that never mentions one:

- the path is the plan, and where the mentor's own suggestion disagrees with it the
  path wins, because a person wrote it;
- a knowledge question is a tutoring moment, not an exam the mentor can shortcut: it
  is not told which answer is correct, and the clause says so, so "I do not have it"
  is the honest answer when a hire asks for it;
- completing a step is asked for, never announced;
- an empty AI-enhanced phase is a subject to talk through, and a step that comes out
  of that goes on the hire's own copy, never on the PM's blueprint.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Six things a testing session turned up, five of them the mentor's own judgement and one
of them the reason it had no judgement to make.

The one that matters most: "where am I?" had two plausible answers — the path a person
wrote, and the work pool, metrics and ledger that say what is true right now — and which
one it landed on depended on nothing the hire could see. That is the two-systems problem
moved inside one conversation rather than solved, so the routing is now stated: path
questions to the path, always; the work pool only when the path has nothing open or the
current step is asking for real work; metrics and the ledger answer how it is going and
never what comes next. Mounted only where both exist.

Also:
- A LOCKED item is never agreed to, however directly asked. It said yes to a step the
  hire's own page refuses to open.
- Items are named by the number their page prints and linked, so "want to take the
  check?" arrives as something clickable.
- Part of a step is a line of its checklist, not the step.
- A refusal that came back NOT PROPOSED has no button: it had been telling hires to
  click something that was never rendered.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two causes, one symptom: the mentor kept telling hires to confirm something that was
never on their screen.

**The search budget.** Only search-only hops consume it — a hop that asks for a backend
tool returns immediately — so a model that searched four times before deciding to act
never got to act: it fell through to the forced final answer, which is generated with no
tools, while the persona in front of it still said "offer `add_path_step`". It took that
at face value. The forced turn is now told plainly that it has no tools and can offer
nothing, and the budget is 6 rather than 4 so a thorough answer still leaves room to act.

**The mechanics were never stated.** The persona said what to offer and never that
offering *is* a tool call. It now says so, that a call missing an argument is refused and
shows no button, and that describing a button is not making one.

Also: an item is written as a linked number in a literal shape (`[#3](/onboarding/...)
"title"`) rather than as a description of one — a small model copies a pattern far more
reliably than it interprets an instruction about formatting.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
There is no skip action and there should not be one: a skip is a request the
PM decides, made on the step's own page. The persona now says so, so the mentor
neither promises one nor ignores the ask.

Adds a test that a spent search budget sends the model the no-tools notice and
keeps it out of the conversation, and counts complete_task among the tools a
hire without a path never hears about.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The mentor brings up a step whose checklist is done but that was never finished,
once, with complete_step; and after a wrong answer it teaches first and offers
one refresher step when the gap is bigger than one explanation. The link example
follows the new /onboarding?step= shape.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Replaces "skipping is not yours to offer" with a clause gated on request_skip:
ask why first, help put the reason into a sentence or two a PM can decide on,
send it in the hire's words, and never promise it will be accepted. A step
waiting on that decision is not pushed, and finishing it is not offered without
saying that it withdraws the request.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The numbers are the page's count, not an order: an item opens once what it
comes after is done, and several can open at once. Added steps are placed with
waits_on and unlocks rather than dropped at the end. A ready-to-close step is
answered where-they-are first, never button first and never with a lecture.
Narrowing a question's options down counts as hinting.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The persona still carried the old ramp: a mentor guiding a hire "from
their first day to doing real work", Task 0 as an action, and a closing
line that called everything else "the path to" a first contribution.
Onboarding is the path a PM's blueprint prescribes; the buddy tutors
along it.

- Identity reworded around the onboarding, not a first contribution.
- Path clauses come first; the arrival clause follows as setup, not
  as part of the onboarding.
- Routing: "how is my onboarding going" goes to the path; the metrics
  answer how their work is going and never how far along they are.
- claim_task_zero removed from the action tools.

Refs #311

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A hire with issues sitting in the pool asked their buddy whether there was
anything they could do and was told there was not. The routing clause is why: it
said the suggested work came up "only when the path has nothing open", so a
half-finished path read as a closed door. That is a gate nobody designed, in the
one place the hire cannot see it -- and picking work up is not a reward for
getting far enough along a curriculum.

So the pool is its own road now, open from day one. An explicit ask for something
to pick up goes straight to the suggested work, "what should I work on" is
treated as the genuinely ambiguous question it is and gets both answers, and
questions about the onboarding still belong to the path.

Two clauses follow the hire past the claim, which is where the mentor used to go
quiet: the task packet and the docs for the work itself, and -- for the path --
reading what a step points at rather than only pointing at it. A step that says
"read issue 123" names something the corpus holds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Keeps both sides: the path tutor from #311 and what dev brought in
(team-mode buddy, board checklist actions, PATH_STEP card, the reworked
onboarding journey). Task 0, the ramp and the path-to-first-contribution
card stay retired; dev code that still reached for them was adapted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Dev's mode tests pinned the hire persona by its old opening line, "the
mentor who guides a new hire". #311 rewrote that line to the tutor along
the path, so the default-mode tests failed and the team-mode negative
check passed for the wrong reason. They now look for "the tutor who
guides a new hire".

Refs #311

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The path read now says where the hire stands on a path whose phases run
side by side. The persona says it too, so the rule holds before the tool
is read and when the hire names a phase the read has not seen them start:
which open phase comes next is the hire's choice, never the lowest number.

Refs #311

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Testing `open_orientation` from the buddy failed on a local model: the
packet came back as invalid JSON, and orientation gave up on the first
parse error, so the hire got no packet at all. Phase assembly already
asks the model once to correct its own JSON; orientation now does the
same before it skips.

Refs #311

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Same failure as orientation, in two more places the buddy reaches:

- Grading a short-text answer: an unreadable reply marked every answer
  in it "could not be graded", which the backend records as wrong. A
  hire with the right answer was told it was wrong.
- A diagram card: one broken quote cost the whole diagram.

Both now ask the model once to correct its own JSON before giving up,
the way phase assembly and orientation do.

Refs #311

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Grading no longer builds a correction prompt it will never send, the
diagram retry uses a _correction_prompt helper like orientation and phase
assembly, and the question clause tells the mentor to start from the
answer the path read now shows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
"Offer flag_to_pm as the last resort" was the only thing the persona said
about escalating, and the mentor applied it to the hire too: an explicit
"flag this to my PM" was refused as not being a PM matter. A new clause,
mounted with the tool, says their asking is the reason. The add-step
clause gets the same exception for a step they ask for themselves.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@DavidLeuter

Copy link
Copy Markdown
Contributor Author

Follow-up fix (8f6adc3): offer flag_to_pm whenever the hire asks for it

While testing, the buddy refused a direct "flag this to my PM". The only thing the persona said about escalating was "offer flag_to_pm as the last resort", and the model applied that to the hire's own requests too.

Changes:

  • New _FLAG_ON_REQUEST_CLAUSE, mounted only when flag_to_pm is available. When the hire asks to flag, raise or pass something to their PM, the buddy offers it in that reply. It never decides for them that it isn't a PM matter and never sends them to the docs first.
  • The add-step clause gets the same exception: a step the hire asks for is offered, because it goes on their own copy of the path.
  • New test in test_buddy_persona.py.

The matching tool-description change is in SprintStartProject/sprintstart-backend#261 (00203027). Verified manually. ruff, format and pyright are clean; 999 tests pass.

@daniilperkin
daniilperkin added this pull request to stack #212 September 27, 2026 19:27
Comment thread src/api/routes/grading.py
Comment on lines +81 to +82
raw = llm.generate(messages)
except LLMUnavailableError as exc:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The retry only handles JSON/Pydantic parsing failures, but _Payload.results currently defaults to an empty list. Therefore a syntactically valid but incomplete response such as {}, or one that omits one requested answer ID, reaches this return without retrying. The caller then reports every missing grade as incorrect even though the model never graded it.

Please make results required and validate that the response contains exactly one result for every requested ID before returning. Missing, duplicate, or unexpected IDs should raise ValueError so the existing correction round is used. For example:

payload = _Payload.model_validate_json(extract_json_object(raw))
graded_by_id = {item.id: item for item in payload.results}
expected_ids = {item.id for item in to_grade}

if (
    len(payload.results) != len(to_grade)
    or set(graded_by_id) != expected_ids
):
    raise ValueError("grading response does not match requested answer IDs")

return graded_by_id

Also change _Payload to:

class _Payload(BaseModel):
    results: list[_GradedItem]

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants