Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 6 additions & 5 deletions deps.edn
Original file line number Diff line number Diff line change
Expand Up @@ -23,8 +23,8 @@
;; (RFC-013, karamazov-3cll.10).
jolt-lang/http-client
{:git/url "https://github.com/jolt-lang/http-client"
:git/tag "v0.0.8"
:git/sha "ccce992d6e3d0035a5ffd1d4364cdb39df4af2f0"}
:git/tag "v0.0.17"
:git/sha "77d7e310a1aab2c7d5ecc9f6cc10f77b7a7f23ba"}

;; the HTTP server: HTTP/1.1 over BSD sockets through jolt.ffi, with a
;; worker pool (a long request no longer holds /health) and channel
Expand All @@ -35,8 +35,8 @@
;; and :port 0 picks a free port.
jolt-lang/ring-chez-adapter
{:git/url "https://github.com/jolt-lang/ring-chez-adapter"
:git/tag "v0.7.8"
:git/sha "124a7399641e409a52a73d2eaa072f031ebbbe38"}
:git/tag "v0.7.10"
:git/sha "8c93fa6fadad917f488358392d496752759a53e6"}

;; the HTTP server's routing: a table of route maps, best-match (a
;; literal segment beats a parameter), no dependencies of its own.
Expand Down Expand Up @@ -66,7 +66,8 @@
;; reconciles the duplicate libssl/libcrypto natives to a single load.
jolt-lang/jolt-crypto
{:git/url "https://github.com/jolt-lang/jolt-crypto"
:git/sha "584a4d32094f0a72613d028deb944de00cd9fc11"}
:git/tag "v0.0.9"
:git/sha "bd19a06c81c92911dbc637b865772c15e1204b05"}

org.clojure/data.json
{:git/url "https://github.com/clojure/data.json"
Expand Down
81 changes: 78 additions & 3 deletions resources/gates.edn
Original file line number Diff line number Diff line change
Expand Up @@ -1166,7 +1166,18 @@
deliberate choice rather than a mistake."}

:exam-ratchet
{:value {:mode :detect :min-drop 2}
{:value {:mode :detect :min-drop 2
;; THE RATCHET AT DONE (karamazov-fgsb, owner's decision
;; 2026-09-29): every test that existed when the run started and
;; is not in the tree unchanged must be explained in done's
;; `changed_tests`, or done is refused. :explain or :off.
:at-done :explain
;; The top-level forms that are tests, compared whole: deftest, and
;; def for the tables tests read (a deleted row is a weaker test).
:test-def-heads ["deftest" "defspec" "def"]
;; Files read as forms; any other test file is compared by its
;; assertion lines (wordlists :exam :assertion-forms).
:clojure-exts [".clj" ".cljc" ".cljs"]}
:kind :policy :capability-tunable? false
:provenance ["karamazov-fgsb"]
:doc "WHAT HAPPENS WHEN A WRITE TAKES ASSERTIONS OUT OF A TEST FILE.
Expand Down Expand Up @@ -3310,6 +3321,23 @@
"{\"name\": \"done\", \"args\": {\"answer\":"
" \"<what you built, and the evidence for it>\"}}\n"
"```"))}
{:name :checklist
;; karamazov-dsfx: an answer silent on a requirement shipped, because
;; no rung knew what the requirements WERE. The checklist is data that
;; exists before the work (acceptance criteria, the held task's tests,
;; what plan declared, the problem's list items — ship.clj decides
;; which apply, :ship-checklist says how), and every item needs an
;; entry by id: met, not met or n/a, with a reason. A not-met ships —
;; an honest limit is a result (karamazov-ylte.1); silence does not.
:when (seq unaccounted-checklist)
:message-form (checklist/refusal unaccounted-checklist)
:provenance ["karamazov-dsfx"]}
{:name :exam-ratchet
;; karamazov-fgsb: a run may change or delete a test that was there
;; when it started, and must say why (:exam-ratchet :at-done).
:when (seq unexplained-tests)
:message-form (exam/refusal unexplained-tests)
:provenance ["karamazov-fgsb"]}
{:name :figure-coverage
:when (and (seq evidence) (seq uncovered-numbers))
;; Says what COVERS a figure, not only that these are uncovered. Run
Expand Down Expand Up @@ -3398,6 +3426,38 @@
:doc "The lexical ship rungs (drg-4026 #44) — adding a rung is a data
edit; the evidence computation stays in ship.clj."}

:ship-checklist
{:value
{;; A list item: "- x", "* x", "+ x", "1. x", "2) x". The first group is
;; the item's text. Lines inside a ``` fence are skipped.
:list-item-regex "^\\s{0,3}(?:[-*+]|\\d{1,2}[.)])\\s+(\\S.*)$"
;; The words an entry's status may be, lower-cased, runs of space and _
;; folded to one space.
:statuses {:met ["met" "done" "yes"]
:not-met ["not met" "not-met" "unmet" "no" "partial" "partly met" "blocked"]
:n-a ["n/a" "na" "n-a" "not applicable" "out of scope"]}
;; How each status reads on the shipped answer's checklist.
:labels {:met "met" :not-met "NOT MET" :n-a "n/a"}
;; The sources only a branch holding NO task answers for: the run's own
;; criteria. A board piece answers for its task (its contract's list
;; items, its tests) and its plan.
:run-level #{:acceptance}}
:kind :policy :capability-tunable? false
:provenance ["karamazov-dsfx"]
:doc "THE SHIP CHECKLIST (samizdat.agent.checklist). `done` refuses an
answer that does not account for every checklist item by id. The
items: the operator's :run :acceptance criteria (a1..), the held
task's :tests when it is prose, not a path (t1), what the branch
declared with plan's `checklist` (c1.., a re-plan adds and never
drops), and the list items of what the branch was asked (p1..: the
run's problem, or a board piece's task contract — a contract with no
list items is itself p1). Acceptance items bind only a branch that
works no task (:run-level); a board revision branch works the task
it carries. An
advisory branch owes none. Each entry is met / not met / n/a with a
reason; the accounted list is appended to the shipped answer so the
critic reads the claims."}

:give-up-reason-floor
{:value 60 :kind :threshold :capability-tunable? false
:provenance ["karamazov-ylte.1" "run dbe64eea-successor"]
Expand Down Expand Up @@ -4279,13 +4339,28 @@
:goal {:type "string"
:description "One or two sentences: what the change does and how it meets the whole ask."}
:rfc {:type "string"
:description "The RFC document in markdown, when this step asked for one."}}
:description "The RFC document in markdown, when this step asked for one."}
:checklist {:type "array" :items {:type "string"}
:description "The requirements this change commits to, one sentence each; done accounts for each by id (c1, c2, ...)."}}
:required ["files" "goal"]}}
"done" {:name "done"
:description "Finish the task and return the final answer."
:parameters {:type "object"
:properties {:answer {:type "string"
:description "The final answer, or the best partial result so far."}}
:description "The final answer, or the best partial result so far."}
:checklist {:type "array"
:items {:type "object"
:properties {:item {:type "string" :description "The requirement's id: p1, a1, t1, c1, ..."}
:status {:type "string" :enum ["met" "not_met" "n/a"]}
:evidence {:type "string" :description "What shows it, or why it is not met."}}
:required ["item" "status" "evidence"]}
:description "One entry per requirement owed, by id."}
:changed_tests {:type "array"
:items {:type "object"
:properties {:test {:type "string" :description "The test's name, or the file's path for a file that is not Clojure."}
:reason {:type "string" :description "Why it changed or went, and what pins the behaviour now."}}
:required ["test" "reason"]}
:description "One entry per pre-existing test this change altered or deleted."}}
:required ["answer"]}}
"give_up" {:name "give_up"
:description "Abandon the task, stating why it cannot be finished."
Expand Down
8 changes: 8 additions & 0 deletions resources/manual.edn
Original file line number Diff line number Diff line change
Expand Up @@ -440,6 +440,14 @@
:summary "Run criteria with an injected shell and judge; one {:name :kind :passed? :output} per criterion, in order. :passed? nil is undecided (not run, or a judge with no verdict) and does not block — fail-open like every judge here. `done` runs the :check criteria (model-free, like the rest of that gate); :feature/verify runs both kinds as Gate 2's other half. Results are journalled per criterion under :acceptance with :at done|verify."}
{:name samizdat.agent.acceptance/refusal
:summary "What the branch reads when criteria failed: each by name with the failure's own words, and what was met. prompts/acceptance-failed.md."}
{:name samizdat.agent.checklist/items
:summary "The ship checklist (karamazov-dsfx): what a `done` answer must account for, item by item, as [{:id :text :source}] — a1.. the acceptance criteria (a branch holding no task only), t1 the held task's :tests when it is prose (its first paragraph), c1.. what the branch declared with plan's `checklist` (a re-plan adds, never drops: state/declare-checklist), p1.. the list items of what the branch was asked (the run's problem, or a board piece's task contract — a contract with no list items is itself p1; fenced code skipped). A board revision branch works the task it carries, so it owes that task and not the run's criteria. The :checklist rung in gates.edn :ship-gates refuses an answer missing an entry; policy is gates.edn :ship-checklist."}
{:name samizdat.agent.checklist/entries
:summary "done's `checklist` argument as {id {:status :evidence}}: a vector of {item, status, evidence} or a map by id, or either as JSON. Status met / not_met / n/a (synonyms in :ship-checklist :statuses); an entry needs a known status and a non-blank reason or it counts as silence (`unaccounted`). A not_met ships — an honest limit is a result. The accounted list is appended to the shipped answer (prompts/checklist-answer.md), its evidence read by the figure rung, and journalled under :checklist."}
{:name samizdat.agent.exam/touched
:summary "The exam ratchet at done (karamazov-fgsb): every test that existed at the run's git baseline and is not in the tree unchanged, as [{:path :test :kind :deleted|:changed :after-hash}]. Clojure test files are compared as whole top-level forms (gates.edn :exam-ratchet :test-def-heads — deftest, defspec, def), reader gensyms folded, and a test moved to another file is not touched; any other file by its assertion lines. The :exam-ratchet ship rung refuses done until done's `changed_tests` gives a reason for each (exam/explanations); a reason journalled earlier in the run (:tests-explained) for the same state of the test counts. The reasons are appended to the shipped answer."}
{:name samizdat.agent.gitdiff/file-at
:summary "A file's content in the run's baseline commit (untracked files included), or nil when it was not there: the tree as the run found it."}
{:name samizdat.agent.judge/parse-yesno
:summary "true / false / nil from a narrow judge reply — the first word of the first line, else of the last. The shape prompts/acceptance-judge.md asks for; change both together."}
{:name samizdat.agent.judge/parse-criteria
Expand Down
5 changes: 5 additions & 0 deletions resources/prompts/checklist-answer.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@


Checklist:
{% for r in rows %}- {{r.id}} [{{r.label}}] {{r.text}} — {{r.evidence}}
{% endfor %}
5 changes: 5 additions & 0 deletions resources/prompts/checklist-missing.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
This answer does not account for every item on the checklist. The checklist is what this work was asked to deliver, fixed before it started, and an item the answer is silent on reads as done when nobody said so. Account for each one below by its id:

{% for i in missing %}- **{{i.id}}** {{i.text}}
{% endfor %}
Call `done` again with your answer and a `checklist` entry for every item: `"checklist": [{"item": "{{missing.0.id}}", "status": "met", "evidence": "what shows it: the test, the command and what it printed"}, …]`. The status is `met`, `not_met` or `n/a`. An item you could not finish is `not_met` with the reason — that ships, and it is the honest answer. Leaving it out does not.
4 changes: 3 additions & 1 deletion resources/prompts/plan-tool.md
Original file line number Diff line number Diff line change
@@ -1,4 +1,6 @@
{% if not-a-path %}plan refused: "{{not-a-path}}" is not a file path. Every entry in files and tests is a bare relative path — test/flight/ghost_test.clj, not the path with a description after it — because a declared file is what you are held to when you finish, and a sentence can never be written. Put descriptions in goal and call plan again with paths only.{% endif %}{% if needs-files %}plan needs at least one file: {"files": ["src/…"], "tests": ["test/…"], "goal": "…"}. Naming a file is the point — it is your hypothesis about where the problem is, and it is what you will be held to when you finish.{% endif %}{% if declared %}{% if planning %}Plan recorded{% if goal %} — {{goal}}{% endif %}. It names: {{files}}.
{% if not-a-path %}plan refused: "{{not-a-path}}" is not a file path. Every entry in files and tests is a bare relative path — test/flight/ghost_test.clj, not the path with a description after it — because a declared file is what you are held to when you finish, and a sentence can never be written. Put descriptions in goal and call plan again with paths only.{% endif %}{% if needs-files %}plan needs at least one file: {"files": ["src/…"], "tests": ["test/…"], "goal": "…"}. Naming a file is the point — it is your hypothesis about where the problem is, and it is what you will be held to when you finish.{% endif %}{% if declared %}{% if checklist %}Checklist recorded, {{checklist}} in all (c1…): `done` must carry an entry for each one.

{% endif %}{% if planning %}Plan recorded{% if goal %} — {{goal}}{% endif %}. It names: {{files}}.

This planning step is complete. A reviewer reads the plan against the requirement next; if it is sent back, you will be asked to revise it. Nothing further is needed from you here.{% else %}Plan recorded{% if goal %} — {{goal}}{% endif %}. You will land: {{files}}.

Expand Down
16 changes: 13 additions & 3 deletions resources/prompts/system-tools.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,22 +10,32 @@ branch_theses({theses})
Propose up to 4 competing plans. The first commits this branch; the rest
become sibling branches that explore independently and share your failure
log, so none of you repeats another's dead end.
done({answer})
done({answer, checklist?, changed_tests?})
Ship. `answer` is REQUIRED and is the run's actual output — the text a
person reads to learn what you did and why they should believe it. A
`done` with no answer is refused and costs you the turn.
Also refused if the answer states figures nothing in the evidence
supports, or engages nothing the problem asked.
`checklist` accounts for each requirement you owe, by its id:
[{"item": "p1", "status": "met" | "not_met" | "n/a", "evidence": "…"}].
The requirements are the list items in the problem (p1…), the operator's
acceptance criteria (a1…), your task's tests (t1) and what you declared
with plan (c1…). A requirement left out is refused; one you could not
meet ships as not_met with the reason.
`changed_tests` says why each test that existed when the run started was
changed or deleted: [{"test": "blank-titles", "reason": "…"}]. A test
changed without a reason is refused.
give_up({reason})
Stop working this line and say why.
```

### Developing at the REPL

```
plan({files, tests?, goal?, rfc?})
plan({files, tests?, goal?, rfc?, checklist?})
Say which files you are about to create or edit, which tests you will
write, and why — one line. When the design step asks for an RFC, `rfc`
write, and why — one line. `checklist` lists the requirements you are
committing to, one sentence each; `done` will ask you for each by id. When the design step asks for an RFC, `rfc`
carries the whole document; it is what the reviewer reads. Every entry in files and tests is a bare
relative path such as test/flight/ghost_test.clj, nothing else: a path
with a description after it is refused, because a declared file is what
Expand Down
5 changes: 5 additions & 0 deletions resources/prompts/tests-explained.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@


Tests changed from the run's start:
{% for r in rows %}- {{r.key}} ({{r.path}}, {{r.kind}}) — {{r.reason}}
{% endfor %}
5 changes: 5 additions & 0 deletions resources/prompts/tests-unexplained.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
This change alters tests that were there when the run started, and the answer does not say why. A test is the definition of done that existed before this work; changing it to get green and changing it because it was wrong look the same in a diff, and only you can say which it was.

{% for t in missing %}- **{{t.key}}** ({{t.path}}) — {{t.kind}}
{% endfor %}
If a change was not meant, put the test back. Otherwise call `done` again with a `changed_tests` entry for each: `"changed_tests": [{"test": "{{missing.0.key}}", "reason": "what was wrong with the old test, or where it went, and what now pins the behaviour"}, …]`.
4 changes: 4 additions & 0 deletions resources/userspace.edn
Original file line number Diff line number Diff line change
Expand Up @@ -59,6 +59,10 @@
{:acceptance-failed "prompts/acceptance-failed.md"
:abandoned-reason "prompts/abandoned-reason.md"
:acceptance-judge "prompts/acceptance-judge.md"
:checklist-answer "prompts/checklist-answer.md"
:checklist-missing "prompts/checklist-missing.md"
:tests-explained "prompts/tests-explained.md"
:tests-unexplained "prompts/tests-unexplained.md"
:adopt-tool "prompts/adopt-tool.md"
:battery-tool "prompts/battery-tool.md"
:heldout-decided "prompts/heldout-decided.md"
Expand Down
Loading
Loading