Skip to content

Latest commit

 

History

1,173 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

murakumo 叢雲

Join the inference network (macOS / Linux)

Install the node CLI with Node.js 22+ and npm already available:

curl -fsSL https://murakumo.cloud/install.sh | sh
export PATH="$HOME/.local/bin:$PATH"
murakumo node init

The installer verifies the release file checksums and installs a pinned runtime under ~/.local/share/murakumo-cli, with a launcher in ~/.local/bin. It needs no sudo, Git, Java or globally installed nbb. It does not download a model, register a node, change your shell profile or start a background service. You can inspect the installer before running it. Set MURAKUMO_INSTALL_DIR and MURAKUMO_BIN_DIR to choose other destinations.

Start your own OpenAI-compatible model server (for example, Ollama or llama.cpp). Use the exact model ID returned by its /v1/models endpoint:

curl -fsS http://127.0.0.1:11434/v1/models
murakumo node doctor --model YOUR_MODEL_ID --local-url http://127.0.0.1:11434/v1
murakumo node check --name my-pc --model YOUR_MODEL_ID --local-url http://127.0.0.1:11434/v1
murakumo node join --name my-pc --model YOUR_MODEL_ID --local-url http://127.0.0.1:11434/v1
  • init creates a private device key on this PC (mode 0600) and leaves an existing identity untouched. Keep ~/.local/share/murakumo-node/identity.json private; it is not a billing-account export. MURAKUMO_NODE_HOME changes this location.
  • doctor checks the local model and identity without writing to the network.
  • check enrolls and sends one signed heartbeat with zero free slots; it never claims jobs. HTTP 201 for that heartbeat proves acceptance, not inference.
  • join serves jobs in the foreground until Ctrl-C. Device sessions renew while it runs. Restart the command after reboot; no persistent service is installed.

Community enrollment starts pending admission. Registration and a fresh heartbeat do not grant AWAI Secure membership, guarantee job placement or prove that a paid request has executed. Contact node support with your node name and public DID if admission is pending; never send the private identity file. Existing operator credentials (MURAKUMO_NODE_CACAO + MURAKUMO_NODE_DID, or MURAKUMO_SERVICE_TOKEN) are still supported. A local server credential goes in MURAKUMO_INFER_LOCAL_TOKEN / VLLM_API_KEY, separately from network authentication.

If diagnosis fails, confirm the server is running and the model ID matches. Missing flags, registration refusal and failed heartbeat checks exit nonzero. The live network view distinguishes connected nodes, ready reports and hosted-server responses. This public CLI packages the existing nbb resident; it is not a claim of native Kotoba runtime migration.

Verified on 2026-09-11

A clean device identity on macOS, an isolated Ollama server and smollm2:135m completed local inference, Community enrollment (201) and authenticated heartbeat acceptance (201). The check did not advertise free slots or fetch the work queue. Community job placement and a paid end-to-end request were not exercised.

Fleet operator tooling

Control plane for the kotoba WASM lattice/mesh across the Mac-mini fleet.

kotoba ships a single-node mesh runtime (kotoba-server with the p2p,realtime-wasm features) and a single-node status command (kotoba lattice ps) — but no fleet-facing control surface. murakumo is that surface: a Kotoba operator (kotoba run / kotoba compile) that runs from your terminal, reaches every node over Tailscale SSH today, installs a resident kotoba mesh node on each, and folds the whole fleet into one view. murakumo.cloud is the Murakumo-native overlay being built to replace the Tailscale/WireGuard dependency with DID/CID identity addressing, policy records, direct QUIC/WebRTC/WebTransport paths, and relay fallback. No agent is installed on the nodes beyond the two kotoba binaries.

GitHub merge governance is also treated as desired state. The exact Actions permissions and app-bound, strict main check are versioned in ADR-260811-version-github-merge-governance.edn and audited against the live API with an administrator-authenticated gh:

kbb --backend sci scripts/check-github-governance.cljk

The EDN is not proof of current enforcement; only a successful live readback is. The decision and blocked-canary evidence are in ADR-260811.

What is each node running? — murakumo fleet ps

One read-only command answers "which model is loaded where, with which context and KV cache, is it busy, and what else shares the node's memory":

murakumo fleet ps                    # every online macOS node
murakumo fleet ps --node joseph      # one node (name or tailnet IP, prefix ok)
murakumo fleet ps --all --json       # every online node, machine-readable
kbb --backend sci scripts/run-task.cljk fleet-ps --node joseph   # from a checkout
joseph (100.82.123.35)  Darwin  16 GB  free 74%  swap 30.6 GB
  model   llama-server pid 32640  2.6 GB  Ternary-Bonsai-2-27B-PTQ1_0.gguf  ctx 32768  kv q8_0/q8_0  fa on  parallel 1  100.82.123.35:8094  slots 0/1 busy
  shares  2.3 GB  kotoba-server  [com.murakumo.kotoba-mesh]
  shares  0.9 GB  node /Users/joseph/.inga/releases/inga-origin-20260911/engine/cl  [com.gftd.inga.reservation.20260911.w4]
  units   com.gftd.inga-node.read-audit.w4, com.gftd.inga.reservation.20260911.w4, com.murakumo.comfyui, com.murakumo.kotoba-mesh, com.murakumo.mishima  (+11 not running)

model rows come from the server's own argv plus a live /slots read (the ctx shown is what the server reports). shares rows are the other processes over 200 MB, each with the launchd / systemd unit that owns it (by pid or parent pid, so a node process started by npm exec under a job is attributed to the job). The probe runs under /bin/sh -s on stdin, never the login shell (the nodes' zsh does not word-split and aborts on unmatched globs), needs no sudo and writes nothing. A node it cannot read is printed as UNREACHABLE with the reason; exit 0 all read, 1 some unreachable, 2 no node matched. Tests: kbb --backend sci scripts/run-task.cljk test-fleet-ps (fixture captured from joseph).

Not to be confused with the etzhayyim murakumo (k3s-on-Lima + Ansible control plane for the religious-corp LangGraph/Pregel cells). This repo is the kotoba WASM mesh layer — libp2p lattice nodes hosting content-addressed WASM components (run / on-http / on-tick / on-kse). Different substrate, on the same hardware.

Qwen3.8 Flash Next Expert-aware NVMe route

qwen3.8-flash-next-ud-iq3-xxs is a single-node, lossless Qwen4Exp route for the 16 GiB M4 mini dan. Planning credits only GGUF bytes actually observed on that node and otherwise requires the complete 81,961,823,936-byte artifact plus an 8 GiB disk floor. The fixed comparison cell keeps the model's top-10, disables cold-expert dropping, temporal prefetch and MTP, and uses four read lanes. The qualified 16 GiB profile keeps the user-space expert cache at zero: the 2,000 MiB comparison cache displaced macOS page cache and reduced the 8-token decode from 1.937 tok/s to 0.306/0.287 tok/s despite a 21.4% hit rate. The 32-token cell confirmed 3.615 tok/s at cache 0 versus 1.473 tok/s for ordinary mmap (2.45x decode; 1.46x including cold load and prefill). Actual token IDs matched across cache 0, two cache-on runs, and mmap. Evidence is in verify/evidence/qwen38-expert-stream-m4-16g-20260829.json.

# leftover JVM infer planner still exists (not operator start; no :infer alias).
# Operator start is kotoba compile + guest instantiateKotoba. There is no kotoba -M.
kotoba compile kotoba/desired.kotoba --target wasm --output target/kotoba/desired.wasm
kotoba compile kotoba/desired.kotoba --target web --output target/kotoba/desired.mjs
# Guest-run: instantiateKotoba on the web artifact (scripts/kotoba-run.sh).
# Release `kotoba run <entry>.kotoba` currently rejects typed forms (CLI source-run gap).

On macOS, Expert-aware reads are real but page-cache bypass is not: neither Linux O_DIRECT nor F_NOCACHE is active in the pinned native runtime. Fleet records therefore expose :page-cache-bypassed? false even if the upstream telemetry prints the requested o_direct=1 flag.

Public API boundary

Clients and applications use https://api.murakumo.cloud exclusively. infer.murakumo.cloud is the authenticated Cloudflare Tunnel origin behind that Worker, not a supported client endpoint. The Worker supplies MURAKUMO_FLEET_ORIGIN_TOKEN; llama.cpp reads the same value from /etc/murakumo/fleet-origin.keys. This keeps model routing, protocol translation, metering, and access policy at one public boundary.

The historical gemma-gad.gftd.ai and gemma-fleet.gftd.ai Tunnel aliases still arrive on gad's loopback port 11434. They must not keep a second copy of the dense 27B model resident. Install deploy/murakumo-legacy-infer-proxy.py as /usr/local/libexec/murakumo-legacy-infer-proxy and enable deploy/murakumo-legacy-infer-proxy.service; it injects the configured fleet origin token and streams requests to the single authenticated server on 8090. Keep llama-server.service disabled while the proxy is active. Rollback is the reverse: stop the proxy, then enable and start llama-server.service.

npm run task -- infer profile [b70|gad|xavier] prints the governed serving profile and exact llama.cpp tuning arguments derived through num → torch → inference → murakumo. This keeps the live unit configuration auditable instead of copying unrelated tuning flags between devices.

The matching resident units live under deploy/systemd/. Xavier has a separate root-owned performance oneshot because nvpmodel -m 0 (MAXN) and jetson_clocks cannot run as the unprivileged model process. The inference unit declares that prerequisite instead of silently serving in 15W desktop mode.

API keys (mk1 capability tokens)

api.murakumo.cloud gates /api/v1/* on a capability token presented as x-api-key or Authorization: Bearer. Two ways to issue one — same implementation (murakumo.apikey), so they cannot drift.

Both need $MURAKUMO_TOKEN_SECRET: the operator's signing secret, the same value the gateway verifies with. Unset ⇒ both refuse, rather than emitting a token that would fail at the gateway with an indistinguishable 401.

CLI

export MURAKUMO_TOKEN_SECRET=…                      # same value as the gateway
kbb --backend sci scripts/run-task.cljk token issue --sub laptop --scope chat --ttl 604800
kbb --backend sci scripts/run-task.cljk token verify mk1.…

The token alone goes to stdout, so it pipes (… token issue | pbcopy); guidance goes to stderr.

MCP

claude mcp add murakumo -- nbb \
  --classpath "src:../org-anthropic-mcp/src" \
  scripts/mcp-server.cljk

Tools: murakumo.issue_api_key (sub / scope / ttl_seconds) and murakumo.verify_api_key (token). Running it grants the client the ability to mint keys within the bounds below, for as long as the secret is in the environment.

Bounds, and why

  • scope must be one the gateway understands (chat | image | all). An unknown scope is refused at issue time instead of surfacing later as a confusing 401.
  • ttl is capped at 90 days (default 30). mk1 is stateless by design — the gateway verifies with no KV or DB round-trip, which also means there is no revocation list. Expiry is the entire revocation story, so an unbounded key would be a permanent one. Re-issue instead of minting long-lived keys.

Compatibility with the verifying gateway is a test, not a claim: kbb --backend sci scripts/run-task.cljk test-apikey asserts that keys minted here are byte-identical to, and accepted by, cloud-murakumo.token.

Historical note: the gateway's 401 used to say kbb -M:murakumo token issue. That command has not been runnable since babashka was retired (ADR-2607173000) — murakumo.core is .clj on babashka.process and bb.edn is gone. The token task above is its nbb port.

The 401 was corrected on 2026-08-13 (ADR-2608133200). It had been changed once already, from kbb -M:murakumo token issue to kbb -M:token issue, but that second command raised ArityException — cloud-murakumo.cli/cmd-token was fixed-arity while the alias passes only issue. Both the arity and the message are fixed now; the message names this repo's token task as the alternative.

What it manages

Operator policy is defined in RULES.md. The liveness/model-placement decision is recorded in docs/adr/ADR-260712-fleet-liveness-model-provisioning.md, Component authority epochs are defined in docs/adr/ADR-260725-component-authority-epochs.md, with recovery commands in docs/FLEET-RECOVERY.md.

Each fleet node runs kotoba-server as a macOS LaunchAgent (RunAtLoad + KeepAlive ⇒ starts at login, self-heals on crash). The node:

  • forms a libp2p gossipsub lattice and advertises a Heartbeat (roles, labels, free-gas, hosted components);
  • hosts WASM components placed by the lattice auction and fires their cron (on-tick), HTTP (on-http), and KSE (on-kse) triggers;
  • persists the components' kqe-assert! output to its kotoba Datom log — i.e. the node is a real PDS/graph writer, not a sandbox.

Quickstart

# Generate the shared operator identity once (kotoba writes the XDG store, mode 0600).
# murakumo reads that same file when MURAKUMO_OPERATOR_SEED is unset. Env still wins if set.
kotoba identity new
export MURAKUMO_KOTOBA_DIR=~/github/com-junkawasaki/orgs/com-junkawasaki/kotoba

kbb --backend sci scripts/run-task.cljk identity      # print the operator DID (never the seed)
kbb --backend sci scripts/run-task.cljk ops status    # fold /health + lattice ps across the fleet
kbb --backend sci scripts/run-task.cljk task run --n 22 --cmd 'hostname'   # fan a batch over the fleet
kbb --backend sci scripts/run-task.cljk token issue --scope chat           # mint a gateway API key

# Operator start (whole-component entries — not kotoba/*_core.kotoba oracles)
kotoba compile kotoba/desired.kotoba --target wasm --output target/kotoba/desired.wasm --json
kotoba compile kotoba/desired.kotoba --target web --output target/kotoba/desired.mjs --json
# Guest-run: instantiateKotoba on the web artifact and call main / run.
# Release `kotoba run <entry>.kotoba` currently rejects typed forms; that is
# a CLI source-run gap, not a host-predicate substitute.
sh scripts/kotoba-compile.sh
sh scripts/kotoba-run.sh

The rest of the control plane is unavailable

This Quickstart used to list fifteen bb … commands. babashka was retired as this workspace's script host by ADR-2607173000; murakumo.core, cloud.clj and overlay.clj are .clj on babashka.process, and bb.edn is gone. The Wave-3 conversion carried over only the shell-out tasks, so sixteen named entrypoints were dropped and have no runnable path today (ADR-2608131600) — nodes, status, provision, mesh, deploy, reconcile, infer, model, dash, cloud, overlay among them. scripts/tasks.edn records which.

Three have been ported to nbb and are the commands above: ops (the nbb port of nodes + status), task (the fleet task plane), and token. Operator start for desired / factory / quic is kotoba run kotoba/<entry>.kotoba (there is no kotoba -M and no :desired / :factory / :quic-* alias). Everything else needs a port — the bodies require babashka.process/cheshire under SCI, so re-registering them is a decision about which of these commands should still exist, not a mechanical conversion.

Every bb … line remaining below in this README is in that dropped set. They are left in place, rather than deleted, because they describe what this repo is for; treat them as a specification of the missing surface, not as instructions. kbb --backend sci scripts/run-task.cljk with no argument lists what actually resolves.

Command surface

command what it does
nodes Tailscale reachability + SSH + whether the mesh binary/agent is present (read-only fleet map)
provision [node|all] rsync the kotoba/kotoba-server binaries, render + load the per-node LaunchAgent (idempotent)
up / down [node|all] launchctl kickstart / bootout the resident mesh node
status [node|all] per-node /health (wasm_executor, peer_count) + lattice ps, folded into one table
mesh [node|all] 2-pass: provision with a fixed P2P port + stable PeerId, collect PeerIds, re-provision with KOTOBA_BOOTSTRAP_PEERS = the others ⇒ ONE gossipsub lattice (fleet-wide auction)
deploy <app.edn> [node] port-forward the node's kotoba port, then kotoba app deploy --publish (clj→WASM → lattice)
reconcile <murakumo.app.edn> [--dry-run|--apply|--watch[=secs]] declarative desired-state (wadm) — fold a fleet manifest vs live placement, report/converge the drift (see below)
kotoba run kotoba/desired.kotoba operator start for signed CID desired-state admission (CID charset, empty-reach, extension chain, publish/pull/reconcile dispatch). kekkai seal, mirror I/O, local apply, and runtime probe are host-listen HOLD
cloud [plan|records|routes|dial|connect <node>|relay <name>|bootstrap] [--cloud=cloud.edn] [--fleet=fleet.edn] plan the murakumo.cloud identity overlay, route hints, driver argv, relay argv, bootstrap order, and control-plane records that replace an external VPN control plane
overlay dial|relay --overlay ... native overlay driver shell: validate canonical dial/relay argv and emit the session record a real stream/packet driver will open
fleet <datom-log.edn> [now-ms] coordination-plane view — fold a kotoba-fleet Datom log into one snapshot (per-work holders · active leases · pending proposals) via kotoba.fleet.view/snapshot. The status of the 20-agent coordination layer, next to the mesh status.
infer probe|plan <model>|provision|up|down|ps|serve|generate distributed inference across the fleet, exo-style — memory-weighted shard plan + pipeline-parallel ring (see below)
task probe|plan|run|report fleet task plane — fan a BATCH of short-lived tasks over the fleet and gather the results (k8s-Job / Ray-tasks shape, next to reconcile's k8s-Deployment shape). kbb --backend sci scripts/run-task.cljk task run --n 22 --cmd 'hostname' (see below)
ops status [all|a,b] fleet read surface, restored on nbb — one ssh round trip per node for /health, wasm_executor, peer links, hosted component CIDs, mesh binary presence and LaunchAgent state. This is the nbb port of the bb-era nodes + status, which have had no runnable entrypoint since the bb.edn removal. kbb --backend sci scripts/run-task.cljk ops status
infer-submit --tokens-file ids.edn give a bare-metal AIUEOS device a real prompt — enqueue one qwen38-generate job (device-P256 worker protocol v3). The device has no tokenizer, so the caller tokenises and the caller detokenises: pass ids, or --tokenize-with torch --vocab-file v.edn --prompt TEXT. Add --wait 120 to poll for the generated token array. Needs the torch sibling on the classpath; it cannot read a .gguf (torch.gguf is .clj, JVM only).

Layout

path role
fleet.edn node inventory (host, roles, labels, port) — the SSoT
connect.edn single connectivity description (read=HTTP-by-CID / live=libp2p multi-transport); node-class → transports
cloud.edn murakumo.cloud overlay declaration: domain, relay set, direct transports, and policy
murakumo.app.edn declarative desired state (wadm manifest): apps × replicas × placement (incl. :reach)
src/murakumo/config.cljk portable config/path/runtime resolution helpers
src/murakumo/component_authority.cljk Component placement/revocation authority: monotonic epochs and exact events consumed by Kototama hosts
src/murakumo/component_authority_store.cljk fsync-backed signed authority outbox: enqueue-before-state, ordered retry, acknowledge-after-delivery
src/murakumo/component_authority_http.cljk HTTPS publisher connecting the durable authority outbox to Kototama’s bounded receiver
src/murakumo/component_authority_deploy.cljk secret-free per-node TLS receiver rollout plan, node audience binding, and overlap-first trusted-key rotation
src/murakumo/connect.cljk connect.edn loader + portable serves-reach? (pure: can a node reach a client class on a plane?)
src/murakumo/cloud.cljk kbb -M:cloud CLI shell: load fleet/cloud declarations and print plans or records
src/murakumo/cloud/plan.cljk portable murakumo.cloud overlay planner: stable IDs, relay choice, node/relay/route/policy records
src/murakumo/overlay.cljk kbb -M:overlay CLI shell for the native overlay driver boundary
src/murakumo/overlay/forward.cljk local TCP forwarder over the sealed relay stream contract
src/murakumo/overlay/dial.cljk host-side dial reachability plus relay hello/frame checks
src/murakumo/overlay/driver.cljk portable overlay driver core: parse/validate canonical dial argv and emit session records
src/murakumo/overlay/relay.cljk host-side relay listener process with identity-aware ack and frame handling
src/murakumo/overlay/runtime.cljk portable overlay runtime adapter registry and execution-report placeholder contract
src/murakumo/deploy/plan.cljk portable deploy/pin helpers: app manifest parsing, kotoba command argv shapes, placement observation, pinned binary copy plans
src/murakumo/identity.cljk portable identity formatting helpers: SHA-256 hex, graph CID, operator token
src/murakumo/persist.cljk portable Datom/atproto repo.write envelope helpers
src/murakumo/ssh.cljk Tailscale-SSH transport (BatchMode, fast-fail, scp/curl-on-node)
src/murakumo/fleet/inventory.cljk portable fleet inventory helpers: node port defaults + selector semantics
src/murakumo/fleet.cljk inventory load + tailscale status enrichment around the .cljc inventory helpers
src/murakumo/provision/plan.cljk portable provision/mesh helpers: p2p ports, bootstrap peers, plist rendering, rsync argv, launch commands
src/murakumo/report.cljk portable CLI report formatting for nodes/status/deploy/reconcile/help
src/murakumo/tunnel.cljk THE transport contract (portable): ssh connection options, the in-band exit-status sentinel every remote command is wrapped in, optional ControlMaster multiplexing, local-forward command shapes
src/murakumo/core.cljk the command implementations + per-node identity derivation
src/murakumo/reconcile/plan.cljk portable wadm planner — PURE desired/observed→plan core
src/murakumo/reconcile.cljk the wadm CLI shell — collect/apply/watch/persist around the .cljc planner
src/murakumo/dash/state.cljk portable dashboard state helpers: snapshot record shape + liveness alert diffs
src/murakumo/dash.cljk snapshotter + web UI + Datom-log persistence around the .cljc state helpers
test/murakumo/cloud_plan_test.cljk offline unit tests for the murakumo.cloud overlay planner
test/murakumo/overlay_driver_test.cljk offline unit tests for the native overlay driver shell core
test/murakumo/reconcile_test.cljk offline unit tests for the pure reconcile core (kbb -M:test)
test/murakumo/smoke_test.cljk namespace-load smoke tests for CLI shell entrypoints
deploy/com.murakumo.kotoba-mesh.plist.tmpl the resident LaunchAgent template
infer.edn distributed-inference config: model registry + head/worker memory policy — the SSoT
src/murakumo/infer/plan.cljk PURE exo-style planner: memory-weighted contiguous layer partition + fits-gate (bb/JVM/cljs/WASM portable)
src/murakumo/infer/moe.cljk PURE mlx-moe planner: single best-memory-node pick, README capacity tiers, expert-ratio verdict heuristic
src/murakumo/infer/engine.cljk PURE engine adapters: plan → llama.cpp --rpc/--tensor-split cmds / mlx.launch ring cmds / mlx-moe serve cmd
src/murakumo/infer/topology.cljk PURE interconnect evidence: folds per-boundary link facts, and gates the :link-gbps choose-strategy is allowed to see (ADR-260815)
src/murakumo/infer/topology_probe.cljk the nbb feed: discover (tailscale + Bonjour vs fleet.edn) / nominal / thunderbolt / measure (receiver-verified transfers)

The authority receiver rollout is deliberately secret-free. Its default artifact source is the pinned sibling checkout at ../kototama; callers may instead pass an explicit release checkout to deployment-plan. Before changing the remote configuration, the rollout verifies that the Kototama runtime is installed at /opt/kototama, and that the node already has a readable PKCS#12 certificate plus a root-managed /etc/kototama/component-authority.secret. Neither TLS material nor passwords are copied by Murakumo. | src/murakumo/infer.cljk | the inference operator: SSH probe → plan → provision → ring up/down → serve/generate (mlx-moe models take the single-node path) | | test/murakumo/infer_test.cljk | offline unit tests for the pure planner/engine (kbb -M:test) | | src/murakumo/task/plan.cljk | PURE task scheduler: eligibility (labels/roles/memory/exclusions) → slot-aware least-filled placement → retry-elsewhere → honest speedup summary | | src/murakumo/task/exec.cljk | nbb execution shell: bounded-concurrency SSH fan-out, per-task timeout, in-band exit-code sentinel, node probing | | src/murakumo/task/worker.cljk | resident remote shells (ssh host bash -s per slot) with per-task framing, timeout-kill + respawn — the transport that removes the per-task ssh cost | | src/murakumo/task.cljk | the task CLI (nbb): probe / plan / run / report + run ledger | | src/murakumo/ops.cljk | the ops status CLI (nbb): the restored fleet read surface, built on the portable dash/state.cljc probe core | | test/murakumo/task_plan_test.cljk | offline unit tests for the pure scheduler (kbb --backend sci scripts/run-task.cljk test-task) | | test/murakumo/task_exec_test.cljk | offline unit tests for the exec shell's pure helpers | | test/murakumo/infer_moe_test.cljk | offline unit tests for the pure mlx-moe planner (kbb -M:test) |

Dashboard + Datom persistence

kbb -M:dash [port=8899] [interval-s=15]   # → http://localhost:8899

dash runs a background snapshotter + a web UI (babashka's built-in http-kit):

  • every interval seconds it polls the fleet (each node's /health, live libp2p LINKS, and the component CIDs it HOSTS = where the lattice placed things);
  • it persists each snapshot to the kotoba Datom log — one atproto.repo.write tx into the murakumo-fleet graph, so the fleet's heartbeat + placement state becomes an append-only as-of history (tamper-evident, queryable);
  • it serves that snapshot as an auto-refreshing page (/) + JSON (/api).

The page shows, per node: HEALTH · WASM-EXEC · LINKS · P2P-PORT · HOSTED (the content-addressed components the lattice placed there) — murakumo status, live, in a browser, with history on the Datom log behind it.

Declarative reconcile — murakumo's wadm

Pull 型の分散 reconcile(Cloudflare 非依存)

従来の reconcile --apply は operator が node ごとに SSH して push する。新しい :desired 経路では、operator は immutable CID を持つ app だけを deterministic に 配置し、Kekkai 共通 envelope(Ed25519 authority、CIDv1、epoch、previous CID)として 複数の独立 mirror に発行する。node は authority を pin して pull し、自分の assignment だけを local Kotoba endpoint へ適用する。適用後は node 自身の鍵で署名した receipt を同じ content-addressed layout へ返す。

# operator start (guest admission). There is no kbb -M:desired.
kotoba compile kotoba/desired.kotoba --target wasm --output target/kotoba/desired.wasm
kotoba run kotoba/desired.kotoba
kotoba run kotoba/desired.kotoba --function run --arg '"publish"'
kotoba run kotoba/desired.kotoba --function run --arg '"reconcile"'

# Native sealed kexe is aarch64-macos only (Linux kexe-verify is HOLD):
#   bin/amu check kotoba/desired.kotoba --jvm-free
#   bin/amu compile kotoba/desired.kotoba --target aarch64-macos --jvm-free --output desired.kexe
#   bin/amu verify desired.kexe

kekkai seal, filesystem mirrors, node-local kotoba app deploy, and the runtime probe are host-listen HOLD in this guest. The leftover Java in src/murakumo/desired_state.cljk still implements that I/O; it is not a start path and has no :desired alias.

Node receipt は deploy command の exit 0 だけでは :applied を名乗らない。app が :probe {:method "POST" :path "/mesh/http/..." :body "{}" :expect-status 200} を desired state に持ち、node-local HTTP probe まで一致した場合だけ :applied になる。 probe のない正常な announce は :accepted、probe 失敗は state/receipt を書かず失敗する。 route 登録だけ成功して component block が無い状態を実行成功として記録しないためである。 常駐化の正本は deploy/com.murakumo.desired-agent.plist.tmpl と deploy/murakumo-desired-agent.sh.tmpl。system LaunchDaemon が 60 秒ごとに local mirror を pull し、再起動後も同じ署名検証・previous-CID・runtime-probe gate を通す。

mirror は filesystem namespace なので、別ディスク、別ホストの mount、object-store mount、閉域同期のいずれでもよい。mirror 自体は信頼しない。同 epoch の異なる CID、 rollback、node が最後に適用した CID を親にしない更新は停止する。artifact 本体は CID で既に local store / mesh / gateway から取得可能である必要があり、この経路は node 上で source build しない。また現段階の apply は既存 reconcile と同じく converge-up のみで、 削除・余剰 replica の自動停止は行わない。

deploy is imperative (compile → distribute → publish, once). reconcile is the declarative half: you write a fleet manifest (murakumo.app.edn) that states what should be running and with what spread, and murakumo continuously folds the live placement against it.

;; murakumo.app.edn
{:murakumo/version 1
 :apps [{:name "kotodama-bot"
         :manifest "../kotoba/examples/kotoba-mesh-app/kotoba.app.edn"  ; clj→wasm
         :cid "bafy…"                          ; once known ⇒ match observed without rebuild
         :replicas 2                            ; desired # of eligible nodes hosting it
         :placement {:labels {:tier "edge" :zone "jp"} :roles ["compute"]}}]}
kbb -M:reconcile murakumo.app.edn --dry-run     # print desired-vs-observed plan, do nothing
kbb -M:reconcile murakumo.app.edn --apply       # re-publish under-replicated apps → auction converges
kbb -M:reconcile murakumo.app.edn --watch=30    # keep it converged; record each plan to the Datom log
kbb -M:reconcile murakumo.app.edn --dry-run --snapshot=snap.edn   # offline, against a recorded snapshot

A plan is per-app: eligible (label/role-matched nodes the auction may place on), running (eligible nodes actually hosting the CID, from dash observation), misplaced (running where not eligible — drift), and an action:

action meaning
:satisfied running == desired, no drift
:place under-replicated → :targets = least-loaded eligible nodes to deploy onto
:over running > desired (reported; murakumo does not auto-evict)
:blocked under-replicated but no eligible node free
:needs-build app has no :cid yet — compile :manifest first

This is the kotoba-mesh ADR's L5 (docs/ADR-kotoba-mesh-wasm-hosting.md): desired state AND observed state are both datoms, so the reconciler is just their diff. The pure core (desired/observed → plan) is unit-tested offline — kbb -M:test, no fleet or SSH needed. --watch writes each plan as a com.murakumo.fleet.reconcile record into the murakumo-fleet graph, so the fleet's desired-vs-observed history is itself a queryable as-of Datom chain (alongside dash's heartbeat snapshots).

Layering vs kotoba's own manifest: a kotoba.app.edn already has per-component :scale + :placement, and kotoba app deploy --publish places them once via the auction. murakumo.app.edn sits above that with the fleet-level desired replica count and the convergence loop (drift detection + self-heal + history) that the one-shot deploy doesn't have. reconcile ≈ wadm; deploy ≈ wash app.

Task plane — task (the fleet as one batch executor)

reconcile converges long-lived apps (k8s Deployment / wadm shape). task fans a batch of short-lived units of work across the same fleet and gathers the results — the k8s-Job / Ray-.remote shape that ADR-2607071400's equivalence table had no entry for (ADR-2607256000).

kbb --backend sci scripts/run-task.cljk task probe                       # cores / RAM / load1 / reachability, per node
kbb --backend sci scripts/run-task.cljk task plan  --n 22 --cmd 'hostname'   # pure placement preview, nothing executed
kbb --backend sci scripts/run-task.cljk task run   --n 22 --cmd 'hostname'   # place → execute over SSH → gather → retry
kbb --backend sci scripts/run-task.cljk task run   --tasks batch.edn         # heterogeneous batch from a file
kbb --backend sci scripts/run-task.cljk task run   --n 8 --labels tier=gpu --cmd './render.sh'
kbb --backend sci scripts/run-task.cljk task report --last 5                 # replay recorded runs from the ledger
  • Placement reuses reconcile's vocabulary — --labels (all must match), --roles (all must be present) — plus --min-mem-gb and --nodes. Unschedulable tasks are reported, never silently dropped.

  • Slots: each node runs at most min(cores, --max-slots) tasks at once (--slots N to override). Placement is least-filled-first (assigned/slots), so a 32-core box takes proportionally more than a 10-core mini.

  • Admission: --max-load / --max-load-per-core hold back nodes that are already saturated by other resident work, and --max-inflight caps fleet-wide concurrency by shrinking slot budgets rather than dropping tasks. Nodes that fail their probe are dropped before the budget is divided, so an unreachable node cannot consume capacity. Every held-back node is listed with its reason.

  • Transport: by default each slot keeps a resident remote shell (ssh host bash -s) and tasks are framed onto its stdin, over a multiplexed connection. Measured on the fleet (110 tasks × 11 nodes):

    transport per-task p50 wall
    one ssh per task 241ms 2572ms
    + connection multiplexing 90ms 1577ms
    + resident workers (default) 15ms 403ms

    --no-worker falls back to one ssh per task (which keeps stderr separate); --no-multiplex also drops connection reuse. In worker mode stderr is merged into stdout (:stderr-merged? true) and a hung task is stopped by killing its worker — the task is reported TIMEOUT, the worker is respawned, and the rest of that node's queue continues. (The multiplex socket path is length-checked: macOS caps a unix socket at 104 bytes and os.tmpdir() alone eats half of it.)

  • probe also reports mesh health — the same round trip curls the node's own kotoba-server /health, so SSH up / MESH down is visible per node.

  • --format edn emits the whole probe/plan/run/report payload as EDN for piping into another tool; p50/p95 per-task latency are reported next to speedup (speedup only compares runs of the same work — latency is the figure that survives a transport change).

  • Retry: a failed task is re-placed on a different node (--attempts, default 2). A task that succeeds on retry is a succeeded task — the run's verdict is the final attempt, with attempts/retried reported separately. This is Ray's task retry, not lineage re-execution.

  • Exit codes are read in band. Measured 2026-07-25: Tailscale SSH on the macOS nodes does not propagate the remote exit status (ssh asher 'exit 7' → 0), while the Linux node returns it correctly. Every command therefore runs in a subshell that echoes a __murakumo_rc= sentinel, which wins over ssh's own code — otherwise 10 of 11 nodes would report silent false successes. This lives in murakumo.tunnel (.cljc), so the bb/JVM control plane (murakumo.ssh → provision / status / deploy / infer / model GC) inherits the same fix instead of each caller re-inventing an && echo ok || echo FAILED workaround. sh-result keeps ssh's own code as :ssh-exit so a transport failure is still distinguishable from a remote non-zero exit.

  • Ledger: every run appends one EDN map to .murakumo-task-ledger.edn (--ledger to move it), the same append-only shape as the relay ledger.

Deliberately NOT provided (see ADR-2607256000 for the honest gap list): a distributed object store / futures, lineage-based re-execution, gang scheduling or placement groups, and an autoscaler. Results come back inline.

Connectivity — connect.edn (one description, every transport)

The mesh has two planes (ADR 90-docs/adr/2606271700-kotoba-transport-planes.md):

  • read — CID-over-HTTP. Transport-untrusted: CID == sha256(dag-cbor(bytes)) is the trust, so any gateway / peer / CDN is equivalent. Browser + edge are first-class here with zero p2p transport.
  • live — libp2p multi-transport (quic native · webrtc browser · webtransport · wss edge). An availability/liveness layer (placement gossip, realtime, 持ち合い).

connect.edn is the single declarative description of this — like a Connect schema, one source of truth spanning every transport. It maps each node class to the transports it speaks:

{:classes {:native  {:read [:http] :live [:quic]                :dialable true}
           :edge    {:read [:http] :live [:wss]                 :dialable true}
           :browser {:read [:http] :live [:webrtc :webtransport] :dialable false}}}

reconcile reads it to honour an app's :placement {:reach […]} — place a component only on nodes that can actually reach its client class:

  • :reach [:browser/read] → any :http node (universal CID pull) → all servers eligible.
  • :reach [:browser/live] → needs a shared live transport with browsers (webrtc/ webtransport). Native nodes speak only :quic today, so such an app reconciles to :blocked — honestly, instead of being placed where browsers can't reach it.

Wiring knob: to make the native fleet serve browser-live peers, add :webrtc to :native :live in connect.edn (one edit) — and wire that transport into the kotoba-net swarm. That single change flips eligibility for every :reach :browser/live app (proven by reach-after-wiring-webrtc-into-native in the tests). QUIC stays the native↔native default; browser/edge are not raw-QUIC peers.

murakumo.cloud — replacing Tailscale/WireGuard

Tailscale and WireGuard are IP/subnet VPN substrates. murakumo.cloud is designed as Murakumo's own identity-addressed overlay instead: nodes are addressed by stable overlay/node CIDs, the control plane is a Datom/atproto graph (murakumo-cloud), and policy is expressed as records rather than ACLs in an external VPN product.

kbb -M:cloud plan        # human summary: overlay CID, nodes, chosen relays, policy count
kbb -M:cloud routes      # route table: direct transport candidates + relay fallback
kbb -M:cloud dial asher  # policy-checked identity-overlay dial hints for one node
kbb -M:cloud connect asher  # canonical murakumo-overlay driver argv for that dial
kbb -M:cloud relay jp-tyo-1 # canonical murakumo-overlay relay argv for one relay
kbb -M:cloud bootstrap      # relays first, then policy-authorized node connects
kbb -M:cloud bootstrap --format=edn # cloud.murakumo.bootstrap manifest for runners
kbb -M:cloud dial asher --from=browser --capability=ssh  # denied unless policy allows it
kbb -M:overlay bootstrap --manifest-file bootstrap.edn   # validate every bootstrap step
kbb -M:overlay run --manifest-file bootstrap.edn         # dry-run ordered overlay runner plan
kbb -M:overlay dispatch --manifest-file bootstrap.edn    # attach runtime adapters to every step
kbb -M:overlay execute --manifest-file bootstrap.edn     # execution-report contract through runtime adapters
kbb -M:overlay adapters                                  # list runtime adapters and implementation status
kbb -M:overlay transports                                # list transport adapters: native relay + external QUIC/WebRTC boundaries
kbb -M:overlay transport-probe --overlay ...             # probe the selected direct transport socket boundary
kbb -M:overlay adapter-plan --overlay ...                # build external QUIC/WebRTC adapter argv + EDN request
kbb -M:overlay adapter-check --overlay ...               # run the configured adapter check command
kbb -M:overlay adapter-supervisor --overlay ...          # plan restart policy for a long-running adapter process
MURAKUMO_QUIC_DRIVER="kbb -M:overlay-adapter" bb overlay adapter-check --overlay ... # use the bundled reference adapter
kotoba run kotoba/quic_driver.kotoba --function run --arg '"check"'   # guest admission
kotoba run kotoba/quic_cert.kotoba --function run --arg '"ensure"'    # host-listen HOLD (BouncyCastle)
kbb -M:quic-cert list                              # show active QUIC material generations/fingerprints
kbb -M:quic-cert rotate --overlay=bafyOverlay --node=bafyNode --host=localhost # rotate active QUIC cert/key
kbb -M:quic-cert verify                            # verify files, fingerprints, and audit hash chain
kbb -M:quic-cert prune --keep=1                    # remove old non-active generations
kbb -M:quic-driver serve --request-edn '{...}'       # QUIC listener; auto-issues cert/key if env is absent
MURAKUMO_QUIC_CERT=cert.pem MURAKUMO_QUIC_KEY=key.pem bb quic-driver serve --request-edn '{...}' # explicit cert/key override
kotoba run kotoba/quic_driver.kotoba   # guest admission; kwik listen is host-listen HOLD
kbb -M:overlay-adapter check --request-edn '{...}'       # reference external adapter driver entrypoint
kbb -M:overlay dial-check --overlay ...                  # probe direct endpoint reachability
kbb -M:overlay dial-check --via=relay --overlay ...      # connect to relay and exchange overlay hello/frame/ack
kbb -M:overlay dial-check --via=relay --frames=a,b,c ... # stream ordered frames through the relay contract
kbb -M:overlay dial-check --auth-key ... --via=relay ... # stream frames with keyed MAC validation
kbb -M:overlay relay-check --overlay ...                 # prove a relay listener can bind locally
kbb -M:overlay serve-relay --auth-key ... --require-auth true --max-frame-bytes 65536 --overlay ... # hardened relay listener
kbb -M:overlay service-plan --listen 127.0.0.1:18022 --service ssh --auth-key ... --via=relay ... # persistent service proxy plan
kbb -M:overlay service-proxy --listen 127.0.0.1:18022 --service ssh --auth-key ... --via=relay ... # persistent byte proxy
kbb -M:overlay local-forward --listen 127.0.0.1:18022 --auth-key ... --via=relay ... # local TCP lines over sealed relay stream
kbb -M:overlay local-forward-bytes --listen 127.0.0.1:18023 --auth-key ... --via=relay ... # local TCP byte chunks over sealed relay stream
MURAKUMO_OVERLAY_AUTH_KEY=... bb cloud bootstrap --format=edn # inject auth-key into driver argv
MURAKUMO_OPERATOR_SEED=... bb cloud bootstrap --format=edn     # derive overlay auth-key if no explicit key is set
kbb -M:cloud records     # EDN records ready to persist/publish into the cloud graph

The current implementation is the deterministic control-plane layer:

  • cloud.edn declares the domain, overlay, direct transports, relay regions, and default-deny capability policy.
  • murakumo.cloud.plan folds fleet.edn + cloud.edn into cloud.murakumo.node, cloud.murakumo.relay, cloud.murakumo.route, and cloud.murakumo.policy records.
  • Relay fallback is deterministic and region-aware (:labels {:zone ...} / :region), while direct paths prefer QUIC/WebRTC/WebTransport.
  • kbb -M:cloud dial <node> is policy-aware: it defaults to from=operator to=fleet capability=ssh, emits direct/relay candidates only when cloud.edn allows that capability, and otherwise returns a policy denial instead of a route.
  • kbb -M:cloud connect <node> turns the authorized dial plan into the canonical murakumo-overlay dial ... argv that a native stream/packet driver can execute.
  • kbb -M:cloud relay <name> turns a relay control record into the canonical murakumo-overlay relay ... argv for starting a relay process.
  • kbb -M:overlay transports exposes the adapter boundary: relay is native today; QUIC/WebRTC/WebTransport are executable external-adapter slots (MURAKUMO_QUIC_DRIVER, MURAKUMO_WEBRTC_DRIVER, MURAKUMO_WEBTRANSPORT_DRIVER) until a JVM/babashka-safe transport is linked.
  • kbb -M:overlay adapter-plan and kbb -M:overlay adapter-check implement the external driver protocol: the configured command receives <action> --request-edn '<murakumo.overlay.adapter-request>', and murakumo records exit code, stdout, stderr, timeout, and missing-adapter failures as EDN.
  • kbb -M:overlay adapter-supervisor produces the long-running process supervision plan for QUIC/WebRTC/WebTransport drivers, including restart policy and max restart count.
  • kbb -M:overlay-adapter is a bundled reference external driver. It implements check, dial, serve, and serve-once against the same EDN request contract, so real murakumo-quic-driver / murakumo-webrtc-driver binaries can be tested against a known protocol shape.
  • murakumo.overlay.quic-driver is the first real external transport driver. It is JVM Clojure using the pure-Java Kwik QUIC stack. It performs QUIC check, dial, serve, and serve-once, opens a QUIC stream, and exchanges the same adapter-hello / adapter-ack records used by the reference adapter.
  • kbb -M:quic-cert issues, lists, and rotates QUIC certificate material under .murakumo/kagi/quic by default (MURAKUMO_KAGI_DIR overrides the path). Files are written owner-only (0600), indexed in index.edn, and tracked by overlay/node/host generation, active generation, fingerprint, and expiry. Issue/rotate/prune operations append to a hash-chained audit log, and kbb -M:quic-cert verify checks both material fingerprints and that audit chain. MURAKUMO_QUIC_CERT and MURAKUMO_QUIC_KEY still override the stored material when supplied.
  • murakumo.overlay.stream models ordered logical streams, so multiple service sessions can share one overlay transport contract.
  • murakumo.overlay.peer keeps deterministic peer discovery/route-selection state from cloud.murakumo.route records.
  • murakumo.overlay.keyring derives per-overlay, per-epoch key material for key rotation while accepting previous/current/next key ids during rollover.
  • kbb -M:overlay service-proxy is the persistent service-proxy entrypoint over the relay byte stream; local-forward* remains the lower-level debug surface.
  • Relay hardening now includes optional auth-required mode and max frame byte limits; rejected frames are not counted as successful dial checks.
  • kbb -M:cloud bootstrap prints the fleet-wide overlay boot sequence: relay processes first, then policy-authorized node dial argv.
  • kbb -M:cloud bootstrap --format=edn emits the same sequence as a cloud.murakumo.bootstrap manifest with explicit phases and executable argv.
  • kbb -M:overlay bootstrap --manifest-file <file> reads that manifest and validates every phase step through the same dial / relay driver contracts before a future runtime opens sockets.
  • kbb -M:overlay run --manifest-file <file> turns a validated bootstrap manifest into a dry-run murakumo.overlay.run-plan, preserving phase order and marking every step as :run or :blocked.
  • kbb -M:overlay dispatch --manifest-file <file> attaches runtime adapter names (murakumo.runtime.quic, murakumo.runtime.relay, etc.) to each runnable step.
  • kbb -M:overlay execute --manifest-file <file> preserves the same ordering and emits a murakumo.overlay.execution-report. murakumo.overlay.runtime owns the adapter registry (relay, quic, webrtc, webtransport, relay-client) and currently returns explicit :would-run execution records; the real socket/relay runtime plugs into this boundary.
  • kbb -M:overlay adapters lists those runtime adapters and their current implementation status.
  • kbb -M:overlay dial-check ... opens a host socket to the planned direct endpoint, proving the dial target is reachable before full QUIC/WebRTC framing exists.
  • kbb -M:overlay dial-check --via=relay ... connects to the relay endpoint and exchanges minimal EDN relay-hello, relay-ack, relay-frame, and relay-frame-ack records carrying overlay, node, principal identity, target, and a small payload.
  • kbb -M:overlay dial-check --via=relay --frames=a,b,c ... streams multiple ordered frames over the same relay connection and verifies one digest-checked ack per frame.
  • --auth-key <secret> on both serve-relay and dial-check adds a keyed frame MAC; relay acks expose :mac-ok? and reject frames with a bad MAC.
  • With --auth-key, relay frame payloads are sealed with AES-GCM on the wire; acks expose :open-ok? / :sealed? after successful decrypt-and-verify.
  • kbb -M:overlay local-forward ... opens a local TCP listener and forwards client input lines as sealed relay stream frames, returning acknowledged payload lines to the local client. This is the first host-side tunnel boundary; byte-stream framing and service proxying can replace the line codec next.
  • kbb -M:overlay local-forward-bytes ... uses the same sealed relay stream but frames raw local TCP bytes as base64url chunks, then decodes acknowledged chunks back to bytes for the local client.
  • cloud.edn declares :overlay/auth-key-env so kbb -M:cloud connect, relay, and bootstrap can inject --auth-key into executable driver argv from the local environment without storing the secret in control-plane records.
  • If no explicit overlay auth key is set, :overlay/auth-key-source :operator-seed derives the driver MAC key from MURAKUMO_OPERATOR_SEED and the overlay CID. The derived key is still only placed in executable argv, not records.
  • kbb -M:overlay relay-check ... opens and closes the relay listener, proving the host can bind the requested port.
  • kbb -M:overlay serve-relay ... starts the minimal host relay process. It accepts TCP connections and returns an identity-aware ack while transport framing is still being implemented.
  • kbb -M:overlay dial ... / kbb -M:overlay relay ... are the repo-local driver shell for those canonical argv. They do not open sockets yet; they validate requests and emit murakumo.overlay.session / murakumo.overlay.relay records so the next implementation layer has a stable executable contract.

The live packet/stream driver is intentionally separate: SSH and provisioning still use Tailscale today, but the CLI now owns the node, relay, policy, and route records needed for a Murakumo-native overlay control plane. The next layer is to bind these route hints to relay processes and host networking so murakumo cli can open the planned identity dials without depending on Tailscale/WireGuard.

Binary pinning (raced-checkout safety)

The shared kotoba checkout is raced by concurrent agents — a sibling rebuild can swap the cli/server protocol out from under a live fleet. So murakumo deploys a pinned binary set it owns under ./bin (the binaries themselves are gitignored; never commit machine-built blobs). What IS tracked is bin/BUILD.edn — the fleet's expected version:

{:source "…/release" :git-sha "d3595506" :version "kotoba 0.1.0" :features "p2p,realtime-wasm"}

So "which kotoba the fleet runs" is auditable in git, while the binary is distributed out-of-band. To bump the fleet version:

# build the new kotoba (cli+server, p2p,realtime-wasm), then:
kbb -M:murakumo pin <its release dir>     # copies binaries → ./bin, rewrites bin/BUILD.edn
git commit bin/BUILD.edn -m "bump fleet kotoba → <sha>"   # the auditable version pin
kbb -M:murakumo provision all             # roll it out

provision/mesh print the pinned version on rollout and refuse if bin/BUILD.edn declares a version but ./bin is empty (clone murakumo → pin the declared sha first).

Distributed inference — infer (the fleet as one model host, exo-style)

The same fleet that hosts WASM components can serve one LLM too large for any single node, exo-style: probe every node's live memory, cut a memory-weighted contiguous layer partition (each node's slice ∝ its usable RAM), and run a pipeline-parallel ring over it. Pipeline parallel is the only scheme that survives the fleet's 1 GbE interconnect (one activation handoff per shard boundary per token — ADR-2605300000); tensor/expert all-to-all stay the Thunderbolt upgrade axis.

The planner + engine adapters are pure cljc (murakumo.infer.plan / .engine) — the same code cuts plans in bb on the operator's terminal, in JVM tests, in cloud-murakumo's CF Worker, and (eventually) inside a kotoba WASM component. Engines:

  • :llamacpp-rpc — every worker runs a small rpc-server (ggml RPC: Metal / CUDA / CPU alike, so macOS minis and linux boxes mix freely); the head runs llama-server --rpc … --tensor-split …, holds the GGUF on ITS disk (weights stream to workers at load, cacheable node-side with -c), and serves the OpenAI-compatible /v1 API. No model download on any worker.
  • :mlx-ring — mlx.launch --backend ring + mlx_lm pipeline sharding for MLX-format checkpoints (all-Apple fleets).
  • :mlx-moe — mu-hashmi/mlx-moe single-node MoE serving: no ring at all. A model registry entry with :model/engine :mlx-moe skips the fleet-wide layer partition — a MoE checkpoint's router only ever activates a handful of experts per token, so mlx-moe loads the model WITHOUT its expert weights, discovers which experts a prompt routes to, and pages only those in from the node's own SSD. One Mac with ≥32 GiB usable memory then serves a checkpoint whose full weights exceed it (mlx-moe's own number: a 46 GB Qwen3-Coder-Next model on a 32 GB Mac). murakumo.infer.moe is the pure single-node planner (best-memory-node pick, README hardware-table capacity tiers, an honest "does this model benefit" verdict from the expert-ratio/shared-expert heuristic); it shapes its plan exactly like murakumo.infer.plan's (one :head? assignment spanning every layer), so infer.engine/commands, infer.credits/settle, and plan/report all work on it unmodified.

For the M4 mini fleet, the efficient GLM-5.2 path is still GGUF + :llamacpp-rpc: 16 GiB minis contribute small contiguous layer shards and cache them with rpc-server -c. mlx-moe is a per-node hot-expert cache for Apple nodes that can hold the non-expert base plus the measured resident expert capacity. The local GLM-5.2 mxfp4 Clojure/Datomic profile is registered as glm-5.2-mxfp4-mlx-moe; it requires a single Apple node at the measured capacity=4/8 tiers, so 16 GiB minis are rejected honestly instead of being counted together as if mlx-moe were distributed.

npm run task -- infer probe                    # live mem/disk/GPU map of the fleet
npm run task -- infer plan glm-5.2-reap50-q2k  # shard plan + go/no-go gate (infer.edn registry)
npm run task -- infer provision                # push rpc-server + raise iogpu.wired_limit_mb
npm run task -- infer ha-provision             # adopt existing binaries into LaunchDaemon + HA key
npm run task -- infer up                       # start the worker ring
npm run task -- infer serve glm-5.2-reap50-q2k ~/models/GLM-5.2-…-00001-of-00004.gguf
npm run task -- infer generate "叢雲とは何ですか"   # OpenAI API → the whole fleet answers

# Hugging Face model cache setup over Tailscale SSH
kbb -M:murakumo model plan trellis-image-large asher
kbb -M:murakumo model setup trellis-image-large        # auto: live non-canary node; asher is fallback
kbb -M:murakumo model status trellis-image-large all
kbb -M:murakumo revive all                             # Wake-on-LAN offline Macs via a live peer

# :model/engine :mlx-moe — same verbs, single-node path (no ring/up/down):
npm run task -- infer plan qwen3-coder-next-mlx-moe    # picks the best-memory node + capacity + verdict
npm run task -- infer provision                        # pip install -U mlx-moe on that node
npm run task -- infer serve qwen3-coder-next-mlx-moe   # nohup `mlx-moe serve` there, OpenAI-compatible /v1
npm run task -- infer generate "叢雲とは何ですか"        # targets whichever host the last plan chose

# Experimental GLM-5.2 mxfp4 hot-expert cache on one 32 GiB+ Apple node:
npm run task -- infer plan glm-5.2-mxfp4-mlx-moe
npm run task -- infer serve glm-5.2-mxfp4-mlx-moe      # capacity/pin-top-k/profile come from infer.edn

The distributed head takes :infer/parallel from infer.edn (currently 2). llama.cpp continuous batching therefore admits two concurrent request slots; the configured 524288-token total context gives each slot the model's full 262144-token training window. Extra requests wait at the head rather than creating unbounded worker processes.

infer provision also installs each macOS RPC worker as the system LaunchDaemon com.murakumo.rpc-worker (RunAtLoad + KeepAlive). On a remote head it creates a dedicated Ed25519 recovery key whose worker-side authorized_keys entry is restricted to exactly one forced command: kickstart that LaunchDaemon. The 30-second head watchdog probes /slots; after the 420-second cold-load grace it first checks for established client connections, because llama.cpp can serialize /slots behind a healthy long decode. A busy head is left alone. An idle head must fail two probes five seconds apart before the watchdog stops it, restarts every RPC worker through that constrained capability, waits until every planned endpoint accepts TCP, and only then starts the head again. ExecStartPre repeats the all-endpoint check, so llama.cpp cannot silently come up with a partial rank set; it also gives the RPC protocol five seconds to settle after TCP starts listening. The watchdog appends machine-readable health/busy/recovery events to /var/lib/murakumo/ring-watchdog-events.log; logrotate retains twelve weekly files. Inspect with systemctl status murakumo-ring-watchdog.timer on the head, tail /var/lib/murakumo/ring-watchdog-events.log, and sudo launchctl print system/com.murakumo.rpc-worker on a Mac worker.

Resident node liveness and work polling

An OpenAI-compatible server that fits a complete model can join the public queue without being mistaken for a live node merely because it enrolled once:

export MURAKUMO_NODE_CACAO=...       # device-issued CACAO; preferred
export MURAKUMO_NODE_DID=did:key:... # must equal that CACAO's issuer
# Existing operator-managed residents may instead use MURAKUMO_SERVICE_TOKEN.
kbb --backend sci scripts/run-task.cljk infer-join \
  --name k16 --model <served-model-id> --local-url http://127.0.0.1:11434/v1

The resident probes <local-url>/models and reports to POST /infer/nodes/<name>/heartbeat every 30 seconds. The server receive time owns liveness; no heartbeat means never, a heartbeat older than two minutes means stale, and only a fresh model probe with a free resident slot means ready?=true. Claiming a job immediately reports zero free slots and finishing it reports the slot free again. MURAKUMO_INFER_LOCAL_TOKEN (or VLLM_API_KEY) authenticates only the local model server and is never reused as the Murakumo control-plane credential.

Is the link fast enough for tensor parallel? — topology_probe

choose-strategy returns "tensor" at 20 Gbps and above, and until 2026-08-15 nothing produced the number it reads — every caller was a test passing a literal, and the production path spelled it (or link-gbps 0), so an unmeasured fleet and a slow one were the same input (ADR-260815).

kbb --backend sci src/murakumo/infer/topology_probe.cljk discover      # nodes from tailscale + Bonjour vs fleet.edn
kbb --backend sci src/murakumo/infer/topology_probe.cljk nominal       # what each NIC claims — cannot lift the gate
kbb --backend sci src/murakumo/infer/topology_probe.cljk thunderbolt   # bridge0 state: the cable-arrived detector
kbb --backend sci src/murakumo/infer/topology_probe.cljk measure --bytes-mib 128   # real transfers, counted at the receiver

The link number is zero unless every rank boundary carries a verified transfer, so no plan the fleet makes today changes; what changes is that pipeline now arrives with :evidence :measured | :partial | :none | :unverified instead of a zero meaning four different things. Measured 2026-08-15: 933 Mbps between minis (a saturated 1 GbE wire — pipeline parallel remains correct), and 27 idle Thunderbolt ports already enrolled in bridge0 on the nine reachable nodes, waiting on cables rather than on software.

Standalone (no RPC ring) — when the model fits on the head alone

The head (this fleet's: an AMD Ryzen AI MAX+ 395 "Strix Halo" APU, Radeon 8060S iGPU) is a real GPU-capable machine, but its :infer/head :bin-dir RPC-ring binary is a CPU-only build — -ngl 999 was a no-op there, so the head's ~40% ring-share ran on CPU alone. For any model whose full weights fit in the head's own memory, skip the ring entirely and run a GPU-backend build standalone instead:

npm run task -- infer down                                              # free the RPC workers, not needed
npm run task -- infer serve-standalone qwen-agentworld-35b-a3b \
  /home/gad/models/Qwen-AgentWorld-35B-A3B-GGUF/Qwen-AgentWorld-35B-A3B-UD-Q4_K_M.gguf

Verified 2026-07-05: 61.5 tok/s vs. 12.7 tok/s for the same model spread across the 7-node CPU RPC ring — real GPU beats network-distributed CPU by ~5x on this hardware. The binary is the official llama.cpp Vulkan release (*-ubuntu-vulkan-x64.tar.gz — Mesa/RADV detects the iGPU cleanly); the equivalent ROCm 7.2 release build detects no device at all (gfx1151/Strix Halo isn't in ROCm's supported list yet). :infer/head :standalone-bin-dir in infer.edn points at the Vulkan build.

Claude Code on the fleet

gftdcojp/local-murakumo's /v1/messages translates the Anthropic Messages API to/from whatever's actually serving here — the same shape of bridge z.ai runs for GLM. tools/claude-murakumo sets ANTHROPIC_BASE_URL/ ANTHROPIC_AUTH_TOKEN/ANTHROPIC_MODEL and launches the real claude binary:

kbb -M:claude                 # or ./tools/claude-murakumo/claude-murakumo

The go/no-go gate is honest about the memory math: plan exits non-zero when the fleet cannot hold the weights (e.g. GLM-5.2-REAP50 MLX-4bit = 214 GiB ✗ on today's 11×16 GiB + head, while GGUF Q2_K = 129.5 GiB ✓) — or, for an :mlx-moe model, when no single node clears mlx-moe's smallest measured hardware tier (32 GiB usable memory). Plans, model registry, and run results (tok/s) are published to cloud-murakumo's /infer/* API.

Identity & no-server-key

kotoba and murakumo share one local operator identity so both CLIs work on one machine without two env seeds.

Generate once with kotoba (kotoba identity new / shared store). That writes

${XDG_DATA_HOME:-$HOME/.local/share}/kotoba/operator.seed

(mode 0600, never committed, never echoed) — the exact kotoba-lang/kotoba#493 contract. When XDG_DATA_HOME is unset that is $HOME/.local/share/kotoba/operator.seed. One path only; murakumo does not also look under ~/.kotoba/.

Resolution: MURAKUMO_OPERATOR_SEED (32-byte hex) if set, else that shared file. Missing both is the existing missing-seed error; murakumo does not invent a seed. identity / report commands print the operator DID only.

kotoba identity new                         # write the shared 0600 seed file
kbb --backend sci scripts/run-task.cljk identity          # same DID kotoba prints
# optional override:
# export MURAKUMO_OPERATOR_SEED=<32-byte hex>

The fleet shares that one operator DID. Per-node identities are derived deterministically (sha256(operator-seed : node-name)) so they are stable and reproducible without storing a secret per node. Autonomous component writes are attributed to the operator (or, with a member CACAO leash, to the consenting member — see com-junkawasaki/kotoba's mesh persistence + etzhayyim's issue-cacao).

Fleet admission — kekkai (zero-trust gate, opt-in)

fleet.edn is the desired inventory; it is not, by itself, an admission record — anyone who can edit fleet.edn and reach a node over Tailscale gets treated as fleet today. murakumo.kekkai closes that gap by gating murakumo.fleet/select — the single choke point every command (nodes, provision, status, mesh, deploy, up/down) resolves its node set through — against kotoba-lang/kekkai, a zero-trust Tailscale-equivalent control plane (coord-LLM proposal ⊣ TailnetGovernor, admission always routed to a human, append-only ledger).

cp kekkai-tailnet.edn.example kekkai-tailnet.edn   # opt in
# edit :status per node ("authorized" | "pending" | "expired" | "revoked")
kbb -M:murakumo nodes    # nodes without :status "authorized" are now excluded,
                      # reported to stderr: "[kekkai] <name>: not authorized (<status>) — excluded from fleet ops"
  • Opt-in, not a breaking default. Absent ./kekkai-tailnet.edn (or $MURAKUMO_KEKKAI_LEDGER), select behaves exactly as before — every command against every fleet.edn node. The gate only activates once that ledger file exists.
  • Deny-by-default. A node not in the ledger at all is treated as "unknown", same as "pending"/"expired"/"revoked" — only an explicit "authorized" entry passes. Being listed in fleet.edn is never, on its own, sufficient (that would make the governor a no-op).
  • Process boundary, not an in-process dep. kekkai rides langgraph/JVM; murakumo's own CLI runs on babashka. Status lookups shell out to the sibling kekkai checkout ($MURAKUMO_KEKKAI_DIR) as a leftover process boundary — not this repo's operator start, and not a :desired/:factory alias. Default operator start is kotoba run.
  • What this does NOT replace. cloud.edn's default-deny capability policy (kbb -M:cloud dial ... capability=ssh) still governs what a given admitted node may reach; kekkai gates fleet membership (is this node operable at all), a layer below that. Real admission ("pending" → "authorized") happens through kekkai's own CoordinationActor elsewhere — this ledger file is the ground-fact snapshot murakumo reads, not the admission flow itself.
  • Pure logic (murakumo.kekkai.gate, env resolution + node partitioning) is unit-tested offline in kbb -M:test; the subprocess shell (murakumo.kekkai) is exercised manually against a real sibling kekkai checkout.
  • Known cost (not yet optimized): one JVM spawn per node. apply-gate shells out to kekkai.cli once per node in the selection (no batching), so enabling the gate on a full fleet still pays one leftover JVM spawn per node (kekkai sibling CLI). That is a named leftover, not murakumo operator start. A launch failure (missing sibling, broken checkout) degrades that node to "unknown" (denied) rather than crashing the command.

Status (honest)

Measured fleet state, 2026-07-25 (kbb --backend sci scripts/run-task.cljk ops status, cross-checked with a task run over all 11 reachable nodes). The claims in this section describe what the CODE does; here is what the FLEET currently looks like:

  • 11 of 12 nodes reachable over ssh (xavier has no ~/.ssh/config entry — see fleet.edn).
  • ~/.murakumo/bin/kotoba-server is installed on 5 of them (simeon, levi, joseph, dan, benjamin) and those 5 answer /health with wasm_executor: ready. The other 6 have no binary — they are un-provisioned, not crashed.
  • Residency IS in place on those 5, and it is a system LaunchDaemon, not a user LaunchAgent: /Library/LaunchDaemons/com.murakumo.kotoba-mesh.plist, RunAtLoad=true, KeepAlive=true, with com.murakumo.kotoba-mesh-watchdog alongside it on dan / simeon / joseph. Every live server reports XPC_SERVICE_NAME=com.murakumo.kotoba-mesh and the full provisioned env (KOTOBA_AGENT_*, KOTOBA_P2P_*, roles/labels/ports, bootstrap peers), so these are launchd-managed services that self-heal as designed. Look in /Library/LaunchDaemons, not ~/Library/LaunchAgents, and note that launchctl list run over SSH shows the user domain only — it will not list this job. (Two earlier passes of this note got that wrong in both directions; the prose above is what was actually measured.) The ai.gftd.murakumo-* agents several nodes carry are a different system — /usr/local/bin/gftd-murakumo … tunnel against murakumo.gftd.ai — and are unrelated to the kotoba mesh.
  • So the gap is provisioning, not residency: the 5 dark Macs (asher, judah, zebulun, naphtali, issachar) have neither the binary nor the daemon. All 5 accept passwordless sudo, so installing the daemon is possible; disk headroom is thin on issachar (1.3G) and naphtali (2.8G). What is still missing is MURAKUMO_OPERATOR_SEED, from which each node's KOTOBA_AGENT_* / KOTOBA_P2P_* identity is deterministically derived — it is not in the environment, not in kagi under that name, and not in the login Keychain under that service.
  • The kotoba-server binaries that ARE installed are byte-identical across all 5 nodes (sha256 f1c813cf…), so redistributing that artifact needs no rebuild.
  • The LINKS column counts peers seen in ~/.murakumo/mesh.log, which survives the binary being removed — read it as history, not as current peering.

Re-provisioning needs the pinned kotoba / kotoba-server binaries (bin/BUILD.edn: kotoba 4f38b74a, features p2p,realtime-wasm,webrtc), which are distributed out of band and are not in this checkout.

  • Works now:
    • wash layer (imperative): provision / resident LaunchDaemon / status-aggregation / per-node + fleet-wide component deploy.
    • cross-node peering (mesh): the 2-pass PeerId-collect + KOTOBA_BOOTSTRAP_PEERS re-provision is implemented — nodes dial each other over Tailscale and form ONE gossipsub lattice, so a component placed anywhere can run anywhere.
    • durable mesh-state projection + dashboard + liveness alerts (dash): heartbeat snapshots persisted to the Datom log, web UI, drift alerts.
    • wadm layer (declarative): reconcile desired-vs-observed plan + --apply convergence + --watch with as-of history. Pure core unit-tested offline (kbb -M:test).
    • kekkai fleet-admission gate (opt-in): select filters to kekkai-"authorized" nodes once kekkai-tailnet.edn is configured.
  • Next:
    • reconcile --apply currently converges up (re-publish under-replicated apps); scale-down / eviction of :over placements is reported but not enacted (kotoba needs a lattice Stop/drain surface murakumo can drive).
    • match-without-rebuild needs the deployed component CID surfaced back into murakumo.app.edn automatically (today you paste :cid after a first deploy).
    • reconcile reads observed placement from node logs (trigger: executed … <cid>); a first-class lattice ps --json on kotoba-server would replace the log grep.
    • :mlx-moe: the planner/engine/CLI path is implemented and unit-tested (kbb -M:test), but not yet run against the live fleet — today's minis are 16 GiB each, below mlx-moe's smallest measured 32 GiB tier, so infer plan qwen3-coder-next-mlx-moe currently reports DOES NOT FIT honestly on this exact hardware until a ≥32 GiB node (fleet or :infer/extra-nodes) joins.

Mishima fast interactive lane

mishima-fast is a primary-only Hermes profile for short interactive turns. It disables project-rule and memory injection, fallback providers, and automatic title generation; output is capped at 256 tokens. The patched Hermes one-shot uses --toolsets none, so it does not construct tool schemas or wait for MCP discovery. It is deliberately not the profile for repository work that needs AGENTS.md, skills, or tools.

The runtime gate uses Hermes's per-request --usage-file receipt instead of the static prompt-size estimate. Cached prefix tokens still count toward the 8,000 token prompt ceiling. A missing receipt, a fallback model, more than one API call, an incomplete turn, a wall time over 90 seconds, or a wrong canary response is a measured failure. The deadline is a per-request safety bound; a small passing canary window does not establish a production percentile SLO.

kbb --backend sci scripts/mishima-hermes-fast.cljk
kbb --backend sci scripts/mishima-hermes-fast.cljk \
  --receipt evidence/mishima-fast.json

License

Apache-2.0.

About

kotoba WASM mesh control plane — resident lattice nodes + fleet ops over Tailscale (bb/clj)

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages