Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
17 commits
Select commit Hold shift + click to select a range
6be1f24
spec: hand lobe on LFM2.5-1.2B (devague /think + /challenge)
OriNachum Aug 10, 2026
8e3ff10
plan: hand lobe on LFM2.5-1.2B — 12 tasks in 4 waves (devague /spec-t…
OriNachum Aug 10, 2026
0ebfc23
feat: hand — the ninth Colleague role on LiquidAI LFM2.5-1.2B (t1-t9)
OriNachum Aug 10, 2026
f4dccfe
docs: hand across the role contract, the counts, and the honesty disc…
OriNachum Aug 10, 2026
f3389ba
fix: the hand lane was missing the cudagraph-estimate off-switch (t10…
OriNachum Aug 10, 2026
7011e54
feat: hand VALIDATED on the Jetson AGX Orin; budget re-derived (t10, …
OriNachum Aug 10, 2026
6437056
plan: resolve r1/r2 against live evidence; record r5/r6 from the Orin…
OriNachum Aug 10, 2026
ec044db
chore: cross-link #181/#182/#183 into the lane, evidence and per-mode…
OriNachum Aug 10, 2026
d74f529
docs: correct two stale claims in my own derivation transcript
OriNachum Aug 10, 2026
68654ea
fix: HAND_ATTENTION_BACKEND was a DEAD KNOB — rendered into .env, rea…
OriNachum Aug 10, 2026
c869147
fix: RETRACT the Orin VALIDATED claim — the budget does not reproduce
OriNachum Aug 10, 2026
ee8d662
feat: hand serves on the DGX Spark — and three boots name why no budg…
OriNachum Aug 10, 2026
fa7dfc5
docs: fold the Spark result and d10 into the delivery summary
OriNachum Aug 10, 2026
7f916f9
test: cover the hand adapter-honesty surface — new-code coverage 51.2…
OriNachum Aug 10, 2026
b44fe32
fix: adapter id collisions, a misleading comment, and build_config co…
OriNachum Aug 10, 2026
3cca67e
fix: CI lint (markdown in my own changelog) and Sonar S8997 monkeypat…
OriNachum Aug 10, 2026
d52b0dc
fix: Sonar S9073 — split composite assertion (0.56.5)
OriNachum Aug 10, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 15 additions & 0 deletions .agex/data/pr/events.jsonl
Original file line number Diff line number Diff line change
Expand Up @@ -96,3 +96,18 @@
{"ts":"2026-07-25T04:45:47.866069+00:00","type":"pr_webhook_posted","pr":157,"event":"pr_replied"}
{"ts":"2026-07-25T04:45:57.235287+00:00","type":"readiness_arrived","pr":157,"waited_secs":0}
{"ts":"2026-07-25T04:46:01.607351+00:00","type":"pr_read","pr":157,"comment_count":6,"threads_unresolved":0,"ci_state":"ok"}
{"ts":"2026-08-10T06:46:40.854177+00:00","type":"pr_opened","pr":184,"title":"feat: hand \u2014 the ninth Colleague role and the fleet's fine-tuning base, on LiquidAI LFM2.5-1.2B (0.56.1)"}
{"ts":"2026-08-10T06:46:41.881031+00:00","type":"pr_review_triggered","pr":184,"command":"/agentic_review"}
{"ts":"2026-08-10T06:50:45.373614+00:00","type":"pr_webhook_posted","pr":184,"event":"pr_opened"}
{"ts":"2026-08-10T06:50:46.870979+00:00","type":"readiness_arrived","pr":184,"waited_secs":0}
{"ts":"2026-08-10T06:50:52.682664+00:00","type":"pr_read","pr":184,"comment_count":3,"threads_unresolved":0,"ci_state":"failure"}
{"ts":"2026-08-10T07:16:51.792749+00:00","type":"readiness_arrived","pr":184,"waited_secs":0}
{"ts":"2026-08-10T07:16:57.343496+00:00","type":"pr_read","pr":184,"comment_count":6,"threads_unresolved":2,"ci_state":"ok"}
{"ts":"2026-08-10T07:27:31.383799+00:00","type":"pr_reply","pr":184,"thread_id":null,"in_reply_to":null}
{"ts":"2026-08-10T07:27:32.459631+00:00","type":"pr_reply","pr":184,"thread_id":null,"in_reply_to":null}
{"ts":"2026-08-10T07:27:32.463044+00:00","type":"pr_batch_replied","pr":184,"count":2,"resolved":0}
{"ts":"2026-08-10T07:27:33.257550+00:00","type":"pr_webhook_posted","pr":184,"event":"pr_replied"}
{"ts":"2026-08-10T07:28:07.155009+00:00","type":"readiness_arrived","pr":184,"waited_secs":0}
{"ts":"2026-08-10T07:28:12.425961+00:00","type":"pr_read","pr":184,"comment_count":8,"threads_unresolved":0,"ci_state":"failure"}
{"ts":"2026-08-10T07:38:52.888182+00:00","type":"readiness_arrived","pr":184,"waited_secs":0}
{"ts":"2026-08-10T07:38:58.375296+00:00","type":"pr_read","pr":184,"comment_count":8,"threads_unresolved":0,"ci_state":"ok"}
2 changes: 1 addition & 1 deletion .devague/current
Original file line number Diff line number Diff line change
@@ -1 +1 @@
unsloth-qat-senses-first-class-orin-variation
hand-lobe-lfm2-5-1-2b
2 changes: 1 addition & 1 deletion .devague/current_plan
Original file line number Diff line number Diff line change
@@ -1 +1 @@
unsloth-qat-senses-first-class-orin-variation
hand-lobe-lfm2-5-1-2b
135 changes: 135 additions & 0 deletions .devague/deliveries/hand-lobe-lfm2-5-1-2b.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,135 @@
{
"plan_slug": "hand-lobe-lfm2-5-1-2b",
"schema_version": 1,
"created": "2026-08-10T03:38:10Z",
"updated": "2026-08-10T06:26:20Z",
"deviations": [
{
"id": "d1",
"what": "t1 also repointed lobes/cli/_commands/route.py and run.py, both outside t1's declared FILES scope",
"task_ref": "t1",
"reason": "Both modules resolved the cheap-tier gear by role_hint=='minor', which no catalog entry carries after the repoint \u2014 they would have raised 'no model with role_hint=minor found in the catalog' at runtime. Leaving them stale to respect the FILES line would have shipped a live regression. route.py's _KNOWN_GEARS and classifier prompt needed the same treatment.",
"affects": [
"t7"
],
"origin": "user",
"status": "approved",
"classification": "acceptable"
},
{
"id": "d2",
"what": "t4 also added a probe surface to lobes/gateway/_readiness.py (probe_backend_adapters, AdapterProbe, ReadinessCache adapter store + current_adapters), outside t4's declared FILES",
"task_ref": "t4",
"reason": "t4's acceptance requires a declared-but-unloaded adapter to be absent from /v1/models AND capabilities. The obvious implementation \u2014 stat the adapter path \u2014 is WRONG here: adapter paths are mounted into the vllm-hand container, not the gateway's, so a filesystem check would false-negative every correctly-configured adapter while still missing the failures that matter (unreadable file, rank above --max-lora-rank, a checkpoint vLLM refused). The honest evidence is the lane's OWN /v1/models, which is a readiness-layer concern and could not live in the four declared gateway files.",
"affects": [
"c19",
"h14"
],
"origin": "user",
"status": "approved",
"classification": "acceptable"
},
{
"id": "d3",
"what": "t6 also added an empty-flag drop rule to lobes/templates/mg-logwrap.sh, outside t6's declared FILES (templates/fleet/{docker-compose.yml,env.example})",
"task_ref": "t6",
"reason": "v1 ships --enable-lora armed with an EMPTY inventory, so --lora-modules=${HAND_LORA_MODULES:-} renders as a bare '--lora-modules=' and vLLM would parse the empty string as a malformed name=path pair, killing the default boot. A compose command list cannot omit an argument conditionally, so the only place to express this is the shared entrypoint. Rule is narrow by construction (matches --flag= exactly) and tested against bare flags, lone --, short flags and non-flag args.",
"affects": [],
"origin": "user",
"status": "approved",
"classification": "acceptable"
},
{
"id": "d4",
"what": "t9 also touched lobes/runtime/_compose.py (GPU_SERVICES) and t3 also touched lobes/profiles/shape_render.py (ROLE_SERVICE), both outside their declared FILES",
"task_ref": "t9",
"reason": "Two role->compose-service maps exist outside the declared file sets and both raise KeyError on an unmapped role: shape_render.ROLE_SERVICE (crashed the goldens regenerator) and _compose.GPU_SERVICES (a shipped-template mirror asserted by test_init_gpu_access). Neither could be deferred \u2014 the goldens do not regenerate without the first, and the second fails CI.",
"affects": [
"t3"
],
"origin": "user",
"status": "approved",
"classification": "acceptable"
},
{
"id": "d5",
"what": "t6's committed lane changed AFTER t6 was complete: VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0 was added when t10's live boot found it missing",
"task_ref": "t6",
"reason": "MEASURED on a physical Orin 2026-08-10: without the knob the lane profiles to 'Available KV cache memory: -9.25 GiB' \u2014 negative, so no boot at any util. Every other lane on the same nightly image already sets it; the hand lane was the sole omission. Nothing about the rendered compose is detectably wrong until an engine profiles memory with it, so no offline test could have caught this at t6 time. Also re-baselined the vllm-hand service hash and the template-defaults golden.",
"affects": [
"t10",
"t8"
],
"origin": "user",
"status": "approved",
"classification": "needs-follow-up"
},
{
"id": "d6",
"what": "t8's committed Orin budget changed AFTER t8 was complete: gpu_mem_util 0.10 -> 0.06, regenerating its goldens a second time",
"task_ref": "t8",
"reason": "This is r1 firing as designed and t10 exercising its authorisation to re-apportion, but it does mean t8's output was not final when t8 closed. The Orin refused 0.10 twice live ('Free memory on device cuda:0 (4.67/61.34 GiB) ... less than desired (0.1, 6.13 GiB)') and boots healthy at 0.06 (KV pool 235,721 tokens, 7.19x concurrency). The reasoning behind 0.10 was plausible and wrong. Net effect: the per-card util mechanism is retained but now carries no divergence \u2014 every card declares 0.06 and only the Orin's is measured.",
"affects": [
"t10",
"c11",
"h8"
],
"origin": "user",
"status": "approved",
"classification": "acceptable"
},
{
"id": "d7",
"what": "t12 delivered ORIN ONLY. Its acceptance says 'Live validation on Thor AND Orin'; Thor is not validated and the plan's scope is therefore not fully met",
"task_ref": "t12",
"reason": "Thor's boot died in LoRA embedding-slot allocation (vllm/lora/layers/vocal_parallel_embedding.py:49, torch.AcceleratorError: CUDA error: device not ready) on a box at 108/122 GiB used with load avg 4.30, and that run predated the cudagraph fix. Three candidate causes remain unseparated: the pre-fix over-reservation, a unified-memory allocation race (the same class as Thor's documented boot-ordering caveat), or an sm_110-specific LoRA-embedding problem. Distinguishing them needs a quiet Thor with the fix in place, which was not available without displacing production containers. Recorded as a scope shortfall rather than presented as a pass: Orin is VALIDATED, Thor and Spark stay DECLARED per #108.",
"affects": [
"c1",
"h1",
"c33",
"h18"
],
"origin": "user",
"status": "approved",
"classification": "needs-follow-up"
},
{
"id": "d8",
"what": "A second dead-knob defect shipped and was caught only by booting the real compose lane: HAND_ATTENTION_BACKEND was rendered into .env by the orin card profile and substituted by nothing",
"task_ref": "t6",
"reason": "builtin/orin.toml declares attention_backend='TRITON_ATTN' for hand, so lobes init renders HAND_ATTENTION_BACKEND=TRITON_ATTN, but the vllm-hand lane never referenced the variable \u2014 the operator reads a configured backend the engine never receives. Same class as d5 (the cudagraph knob) and same detection story: valid compose, green suite, wrong behaviour only visible when an engine runs it. It also explains why my earlier hand-rolled docker-run probe and the committed lane disagreed \u2014 the probe set the backend manually, the lane could not. Fixed via --attention-config (VLLM_ATTENTION_BACKEND is gone on this nightly) plus a test asserting EVERY profile-rendered key is substituted by the template, verified to fail with the fix reverted.",
"affects": [
"t8",
"t12"
],
"origin": "user",
"status": "approved",
"classification": "needs-follow-up"
},
{
"id": "d9",
"what": "RETRACTED the Orin VALIDATED claim before the branch became a PR \u2014 the budget does not reproduce",
"task_ref": "t12",
"reason": "The acceptance transcript was written after ONE successful boot at gpu_mem_util=0.06 (available KV 2.7 GiB, pool 235,721 tokens, 7.19x). Two later boots of the same configuration on the same box profiled 0.14 GiB and 0.09 GiB and refused to start \u2014 vLLM clamps its budget against actual free memory at startup, and this is a SHARED box whose free memory moved ~2.7 GiB across the runs. One success against two refusals is an unstable budget, not a validated one. The functional results survive (they were observed on a real serving engine and do not depend on the budget): the lfm2 parser, the structured tool_calls array, the bf16 sentinel, the unknown-id 404. The budget claim does not. Transcript renamed to ...-partial-... with the retraction at its head; orin.toml, the per-model doc and machine-profiles.md all now read DECLARED. NO card is validated for hand.",
"affects": [
"t8",
"t10"
],
"origin": "user",
"status": "approved",
"classification": "needs-follow-up"
},
{
"id": "d10",
"what": "t12 live validation extends to the DGX Spark: functional PASS on a third card, budget again NOT reproducible (6.21/3.34/3.54 GiB at identical 0.06) \u2014 the three-run spread identifies the MECHANISM (vLLM profiles against free-memory-at-that-instant on unified memory), which retro-explains the Orin retraction as a property rather than a fluke; all four cards stay DECLARED",
"task_ref": "t12",
"reason": "the user asked for a Spark answer; the box's earlier boot failures turned out to be memory exhaustion (swap 100% full, browser holding 31.7 GiB), not a lane defect \u2014 with memory freed the committed lane booted in 71.75s and served correctly",
"affects": [
"t10"
],
"origin": "user",
"status": "approved",
"classification": "needs-follow-up"
}
]
}
Loading
Loading