Skip to content
View noctem-o's full-sized avatar

Highlights

  • Pro

Block or report noctem-o

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
noctem-o/README.md

George Gallagher

AI systems you can inspect, evaluate, adapt and control.

Agent runtimes · evaluation & adaptation · trustworthy infrastructure · local inference

Recension Research Greater Manchester, UK

Founder + engineer at Recension Research

Endophasia · Magpie · Deadbolt · Cogitator


I build experimental infrastructure for AI agents and AI-operated systems: ways to observe runtime behaviour, evaluate it reproducibly, preserve evidence, keep authority outside the model, and verify what happened afterwards.

The recurring question is simple:

How do increasingly capable agent systems remain inspectable, evidential and governable without pretending one mechanism solves everything?

My work keeps five jobs separate:

observe  →  evidence  →  policy  →  permission  →  verify
runtime     support      standing    authority      recompute

Current work

Project What it does Surface
Endophasia Runtime-neutral experimental substrate for coding agents: trajectory capture and comparison, replay, capability studies, governed steering, cross-runtime adapters and controlled evaluation. TypeScript · Node.js · agent harnesses · evals
Magpie Replayable epistemic memory with signed history, typed evidence, provenance, policy-defined standing and detached standing receipts. Rust · SQLite · Ed25519 · portable verification
Deadbolt Permission engine for AI-operated systems. Proposals remain untrusted until a typed route, policy, exact lease and confirmation authorize an effect. Rust · capability security · MCP · rollback
Cogitator Recorder and verifier for agent runs, with canonical events, policy interception, replay, drift diagnosis and tamper-evident witness roots. Rust · BLAKE3 · verification · release assurance

Each project stands alone. Together they explore a division of responsibility:

agent runtime
     │
     ▼
Endophasia
observe · steer · compare
     │
     ├──── Magpie
     │     evidence · provenance · standing
     │
     └──── Deadbolt
           permission · effects · receipts

Cogitator
reproducible run verification

This is a division of responsibility, not a dependency graph. Observation is not authority, and a successful evaluation is not permission to replace a running system.

Recent empirical work

Endophasia has moved from architecture into repeated live agent studies. The current research separates three different claims:

exact replay               did the recorded session reproduce under recorded responses?
fixed-condition repeat     do repeated live runs under the same declared conditions agree?
perturbation invariance    does behaviour hold when an input that should not matter changes?
Study Evidence
Pinned-environment repeatability 240 live trials under a controlled Pi + local-model configuration, with every manipulation check passing.
Path sensitivity 360 trials across 30 working-directory paths, exposing task-dependent behavioural sensitivity to a nominally irrelevant perturbation.
Discriminating-task construction 72-trial screen + 20 fresh confirmation trials, selecting tasks that avoid saturated ~0% / ~100% success regimes for later candidate evaluation.
Completion-cap / test-time compute 192-trial main study + 40-trial pilot: raising the completion cap improved pooled success on the two discriminating tasks from 10/24 to 19/23, while also exposing why evaluation workspaces need stronger isolation.
Governed steering audit Re-verified 364 recorded interventions while hardening consumption semantics, STOP handling and authorization expiry.

The useful outcome is not a prettier benchmark number. It is a better experimental instrument: one that can expose where sampling, environment, control flow and harness policy actually change an agent trajectory.

The runtime surface is also moving beyond one harness: a merged ACP v1 vertical slice has exercised a real OMP session through the standard protocol while keeping unsupported semantics explicit rather than treating protocol compatibility as capability equivalence.

Evaluation-driven adaptation

The next layer is EVOLVE: treating prompts, cognition policies, harness components, tools, memory, models, environments and inference budgets as explicit candidate changes.

observe → propose → isolate → evaluate → compare → admit → promote

Current research directions include:

  • pluggable evaluation, environment, adaptation and optional training providers
  • harness and policy search
  • test-time compute allocation and stopping policies
  • held-out checks and explicit promotion gates
  • trajectory-preserving evidence for candidate comparison

RL is one possible adaptation mechanism. Training can produce a candidate; it does not decide what the evidence establishes, nor whether that candidate may replace anything.

Engineering surface

Languages: Rust · TypeScript/Node.js · Python · Bash/PowerShell
Agent/runtime: Pi · MCP · ACP · OpenAI-compatible APIs · local agent harnesses
Inference: llama.cpp · vLLM · SGLang · TensorRT-LLM · CUDA · GGUF · quantisation · long-context serving
Systems: Linux · Arch · Nix · systemd · containers/virtualisation · networking · low-level debugging
Delivery: GitHub Actions · reproducible environments · conformance testing · golden fixtures · cargo-dist · release verification

I use frontier and local models as a structured engineering workforce: architecture-first decomposition, bounded implementation briefs, maker/checker separation, adversarial review and human-controlled promotion.

Working principles

observation  != inference
evidence     != standing
proposal     != permission
evaluation   != promotion
simulation   != real execution

These projects are experimental research software. I try to keep claims narrow, make failure states visible, test hostile cases, and document what each system does not establish.


Make capability legible before making it larger.

recensionresearch.co.uk

Pinned Loading

  1. endophasia endophasia Public

    A harness-neutral cognition, observation, control and evolution substrate that can attach to agent runtimes without replacing their native execution loops.

    TypeScript

  2. magpie magpie Public

    Local-first, replayable memory for AI systems. Signed history, provenance, deterministic replay, and explicit policies for what evidence is allowed to conclude.

    Rust

  3. deadbolt deadbolt Public

    A local-first permission engine for AI-operated systems. Typed authority, explicit consent, bounded execution, receipts, and rollback.

    Rust

  4. cogitator cogitator Public

    Tamper-evident AI agent audit harness. Cryptographic witness chain, pre-call policy interception, and byte-stable replay. Prove what your agent did, tried to do, and was blocked from doing. Written…

    Rust