Run a tabletop game where some of the seats are AI agents — and find out afterwards how it actually went.
Agents start here → SKILL.md — when to call this, worked examples, MUST/MUST NOTs. Family contract: FAMILY.md.
Three deterministic referees already exist for the rules of such a table: srdcheck for rules verdicts, charactercheck for character sheets, dmcheck for table conduct. They are libraries. The thing that makes them a table — the chat transport, the session file, the game state, and the instrumentation that turns an evening into something you can learn from — is what this repo is.
It is an example, not a framework. Small enough to read in one sitting, and it expects to be edited.
Discord writes are disabled by default. The outbound helper validates and prepares an immutable operation, denies unintended mentions, durably receipts every delivered chunk, and commits only complete delivery. It remains off until the operator explicitly consents and the production gates in docs/TRANSPORT.md are satisfied.
pip install tablekit
tablekit init # writes table.json
python3 examples/demo_session.py # a whole synthetic session + its reportThe table speaks English. Nobody at it learns a syntax.
Text RPGs solved this badly for forty years. Zork and Moria and everything after them handed you a verb list, because a parser in 1980 could not understand a sentence. That vocabulary was never a design choice — it was a workaround for a missing capability, and it is exactly what makes those games feel like operating a computer instead of sitting with someone who is running a world for you.
The capability is no longer missing. So an agent GM that asks people to learn commands has voluntarily reproduced the limitation and thrown away the only thing it has over a parser game.
Concretely, in this repo: there are no chat commands, no !tokens, no
keywords, no "advanced mode". Whatever someone writes in the channel is
dialogue. tests/test_no_player_commands.py enforces that, because a stated
principle drifts and a test does not.
An earlier version of this kit got this wrong — it shipped a six-token feedback
vocabulary. It was cut in 0.2.0 and the reasoning is preserved in
tablekit/uxr.py.
Everything a session produces currently goes into one append-only JSONL
file, which is simultaneously the play ledger and the telemetry stream.
dmcheck can consume the narrow legacy turn/act/event play lane directly;
complete cross-tool handoff requires the explicit TableEvent adapter tracked by
PORT-002 and must never silently discard incompatible rows.
For completed legacy ledgers, the unreleased migration adapter writes a separate canonical stream and a machine-readable coverage summary:
tablekit migrate-events --ledger table-data/session.jsonl \
--campaign campaign-id --session session-id --out session.events.jsonlThe migration is deterministic and always self_attested. It hashes stable
migration event IDs, preserves source-native IDs only as source/correlation
metadata, reports every unprojectable telemetry row, and refuses malformed or
zero-compatible ledgers. It never rewrites the append-only source ledger.
| lane | answers | how |
|---|---|---|
qa |
did the machinery work? | automatic |
qc |
was the refereeing correct? | automatic |
ux |
what did the seat's evening look like? | automatic |
uxr |
how did it feel from a seat? | inferred from ordinary speech, then asked about at the close |
out |
did any of it actually work? | intent → payoff pairs |
Inbound handling is deliberately fail-closed. A source-native message ID is
committed before routing and deduplicates replays, including messages that are
rejected or quarantined. An inbound can resolve at most one typed obligation:
an explicit pair_id must name the compatible open pair. Being the only open
candidate does not grant an uncorrelated event authority to close it. Blank
messages acknowledge nothing; missing, unknown, cross-seat, and cross-kind
correlations remain open and are surfaced in qa.route. This is the
fail-closed bridge until PORT-002/PORT-003 provide host correlation and
identity.
Identity fallback is exact after Unicode/case/whitespace normalization, never a
substring guess (Will does not capture William). Unknown speakers remain
unknown. Relayed rolls additionally require a configured relay name, an
actual bot flag, exact sheet ID or exact configured alias, a numeric result,
a plausible declared natural-die range, and explicit pair correlation before
resolution. Ordinary prose such as 14 or nat 20 is advisory only;
damage, healing, hit points, movement, and other non-roll quantities do not
auto-resolve a roll.
This is the one that does not exist elsewhere, and the one the no-syntax rule shapes most. A transcript of a great session and a transcript of a session everyone endured look about the same — whether a description landed, whether someone was quietly lost, whether a player got to do the thing their character exists for, none of it leaves a trace.
So it is captured two ways, both of them in plain English:
Inferred during play. "Can we get moving" is the signal. The agent GM is already reading every line; classification does not need a new interface, it needs the GM to record what it already understood. Six internal buckets — pacing, floor, comprehension, resonance, spotlight, continuity — never surfaced to anyone at the table. Every inferred signal stores the player's own words, so the inference is auditable by whoever was in the room.
Asked at the close. Two or three questions in your own words, the way any
decent GM already ends a night. tablekit debrief prints them.
The second half is not optional politeness — it is the structural check on the
first. A GM that does not notice friction cannot record friction, so inference
alone inherits exactly the blind spots the lane was built to route around. An
independent model pass over the same transcript (--source local) is the other
check; where it and the GM disagree is itself the interesting part.
tablekit reportRun tablekit qc after the last session mutation first. QC records its input
and config digests, exact checked-through cursor, evaluator coverage/errors,
and explicit finding transitions automatically. A report exits 2 and names the
gap when that run is missing, stale, invalid, or incomplete; an attention-only
run exits 1 consistently in both commands. Findings remain visible even when
the report refuses an aggregate claim.
Defects first, then what the seats said, then whether the craft moves worked, then the shape of the evening, then whether the plumbing held.
## Defects
✗ undeliverable_cue: beat 4: cue addresses agent seat 'Vesh' but does not
contain its literal mention — likely to be dropped in transit
✗ seat_quiet: Vesh has not said anything for 49 minutes and has not been
checked on
## From the seats (5 signals)
Patterns (>= 3 — worth acting on):
- pacing ×3 — this stretch felt slow from that seat; seats: brae, rowan;
beats 5, 6, 6 [dm/local]
## Did it work (outcome pairs)
cue 3/4 good (75%) median 60.0s
roll 0/1 good (n=1 — too few to state a rate; counts shown instead)
Note the ordering: the undeliverable cue is reported before the silence it caused. And note what is absent — there is no score at the bottom, and there is not going to be one.
1. Give players a command syntax. Above. It is a product non-goal, not a backlog item.
2. Let an inference accuse anyone. Signals classified from speech are advisory in every path and can never become a defect. Something a model thought someone meant does not get to be a finding.
3. State a rate from a handful of observations. Below three occurrences, the report lists individual moments with their beat numbers and refuses to describe a tendency. This is a correction of a real error: one player asking what a word meant was once generalised into "unfamiliar diction fails at this table", and their actual position was the opposite.
Findings come in two severities and the difference is load-bearing: defects are boundary crossings (pass/fail, worth interrupting for); attention items are dosage readings (real signal, no boundary crossed).
tablekit/ the Python package — events, config, lanes, report, CLI
transport/ Discord gateway listener (JSON out; knows nothing of the schema)
engine/ optional boardgame.io combat state: initiative, HP, seeded dice
bootstrap/ a first-session kit for an agent about to GM for the first time
examples/ a full synthetic session you can run right now
docs/ INSTRUMENTATION.md is the one to read
Same convention as the sibling projects: 0 coverage-complete clean,
1 findings or advisories worth looking at, 2 refused/incomplete. Exit
2 covers bad input, stale or failed QC, and insufficient evidence; it is
deliberately distinct from "the session was fine." Current evidence remains
self_attested until a host boundary outside the evaluated agent owns it. An
unresolved routing quarantine is exit 1 even when it safely prevented a bad
state transition; operator work still remains.
- docs/QUICKSTART.md — a table running in about ten minutes
- docs/INSTRUMENTATION.md — the lanes, the event schema, and the provenance of every default
- docs/TRANSPORT.md — mandatory mentions, Discord capability and recovery boundaries, the relay tax, and the other things that cost an evening to learn
- docs/DATA_SAFETY.md — paths, permissions, durability, plaintext storage, and the current integrity boundary
- docs/CONTRACTS.md — shared EvaluationResult/TableEvent schemas, versioning, fixtures, and the migration boundary
- docs/THREAT_MODEL.md — principals, capabilities, abuse cases, and why current evidence remains self-attested
- bootstrap/CORE.md — if an agent is about to GM for the first time, this is the page to hand it
tool.json describes the command surface — all of it operator-side, none of it
touched by anyone at the table. tablekit schema prints the repo-local event schema,
typed pair outcomes, signal buckets and exit codes as JSON. The session file is
JSONL with one object per line and a documented type registry; reading it needs
no library.
tablekit contract evaluation, tablekit contract event, and
tablekit contract golden print the packaged cross-suite contracts and fixture.
Treat it as table-private data and as evidence, not a cryptographically trusted
audit log; the exact boundary is documented in
DATA_SAFETY.md.
For outbound use, call post.prepare() first when inspection is enough. A live
post.post() requires transport.write_enabled=true; supply and persist an
operation_id, and treat every status other than committed as not fully
delivered. post.resume() can recover the complete bounded plan from the
ledger after interruption. The exact retry/reconciliation contract is in
TRANSPORT.md.
MIT.