Skip to content
This repository was archived by the owner on Sep 29, 2026. It is now read-only.
This repository was archived by the owner on Sep 29, 2026. It is now read-only.

Train the body adapter around the fixed connectome #5

Description

@TailsProwerWorks

Train the body adapter around the fixed connectome

Summary

The current PersonConnectome controller loads a MaleCNS-derived connectome, but the fly neural circuit is being driven through heuristic human-body mappings. In-game testing shows neural activity and hazard responses, while ordinary locomotion mostly produces back-and-forth wiggling. We need a reproducible training pipeline for the sensory encoder and motor readout so the fixed connectome can control a People Playground human more effectively.

This issue is about training the adapter around the connectome, not blindly changing millions of biological synapses.

Dependency on the modular refactor

This issue depends on issue #4. The modular refactor remains required and is not replaced by this training work. Land the Core/Adapters/Presentation/Composition boundaries first, then add training through those interfaces.

Training must not introduce a second brain/runtime implementation, duplicate the connectome loader, or put optimizer logic into the People Playground entry point. The trainer should modify versioned encoder/readout parameters between episodes while the fixed connectome and game-facing adapter remain explicit and testable.

Before this issue is considered complete, #4 must provide stable contracts for PersonObservation, SensoryFrame, NeuralMotorFeatures, and MotorCommand, with one authoritative RuntimeBrain shared by runtime and tests.

Goals

  • Make stable locomotion the first trained behavior: stand, walk forward, stop, and turn.
  • Preserve the MaleCNS graph and its provenance by default.
  • Learn or tune the mapping from People Playground observations to sensory input populations.
  • Learn or tune the mapping from selected neural outputs to human-body actions.
  • Support hazards and reactive behavior: avoid lava/fire/electricity, recover from disturbances, and stop when incapacitated.
  • Provide a deterministic offline training/evaluation path before using the game runtime.
  • Export a versioned, validated adapter asset that the mod can load at runtime.

Non-goals

  • Do not retrain every connectome synapse as the first approach.
  • Do not claim biological equivalence between a fly brain and a human body.
  • Do not add an external Python, network, or cloud dependency to the shipped mod.
  • Do not make weapon use the first milestone. Gripping, aiming, and trigger control can follow once locomotion is reliable.

Proposed architecture

PersonConnectome.Core
  MaleCNS asset/loader
  LIF simulation
  deterministic stepping
  state/checkpoint support

PersonConnectome.Training
  episode/environment interface
  sensory encoder parameters
  motor readout parameters
  reward functions
  mutation/optimizer strategy
  checkpoints and evaluation reports

PersonConnectome.PeoplePlayground
  human observation adapter
  liquid/status/damage sensing
  action application
  game-time telemetry

PersonConnectome.Runtime
  loads the trained adapter asset
  validates version, graph hash, ranges, and dimensions
  runs the fixed brain plus adapter in-game

The connectome should be immutable during normal training. Training parameters should be explicit and separate, for example:

  • input population IDs and gains
  • input normalization/clamping
  • temporal filters and decay
  • motor population IDs
  • left/right/forward/stop readout weights
  • limb/grip gains
  • action smoothing and dead zones
  • optional recurrent adapter state, with bounded size

Observation and action model

Record a fixed-rate PersonObservation frame containing only data available from the supported game API:

  • position, velocity, facing, angular velocity, and body tilt
  • head/limb positions and joint angles where available
  • contact/support state and collision impulses
  • nearby object direction/distance/material categories
  • light, sound, vibration, and line-of-sight cues
  • liquid identity, wetness, exposure duration, and estimated severity
  • heat/cold, fire/lava, electricity, oxygen, blood, infection, poison, anesthesia, and injury state
  • target direction and optional curriculum objective

Decode a bounded MotorCommand such as:

  • forward/backward locomotion
  • turn left/right
  • stop/brake
  • per-limb target or motor effort
  • grip open/close
  • optional recovery/brace action

Every action must be clamped and rate-limited before reaching game APIs.

Training stages

Stage 0: deterministic harness

  • Add a headless environment with a simplified 2D body model and the same observation/action contracts.
  • Add fixed seeds, replayable episodes, checkpoints, and golden evaluation traces.
  • Verify that identical asset, adapter, seed, and episode produce identical outputs.

Stage 1: locomotion curriculum

Train in increasing difficulty:

  1. remain upright and reduce unwanted oscillation
  2. hold a heading
  3. walk toward a target direction
  4. stop on command
  5. turn toward a target
  6. walk around simple obstacles
  7. recover from small pushes and uneven contact

Keep the connectome fixed. Start by optimizing only decoder weights and gains; add encoder parameters only when the decoder-only baseline cannot solve a stage.

Stage 2: sensory grounding

Compare learned inputs against hand-authored baselines. Train the smallest useful input set first:

  • forward/left/right target direction
  • support/upright error
  • velocity error
  • touch/collision
  • threat/heat/fire/lava

Add light, sound, chemicals, and other effects only after locomotion has a stable baseline. Log which input populations and output populations actually contribute to the behavior.

Stage 3: reactive survival

Add rewards and penalties for hazard avoidance, safe stopping, recovery, and survival. Test water, acid, poison, anesthesia, electricity, fire/lava, damage, blood loss, infection, and unconsciousness independently and in combinations.

Stage 4: limbs and interaction

After locomotion passes its acceptance thresholds, train reaching, gripping, releasing, and simple object interaction. Treat weapon/trigger behavior as a separate opt-in experiment with its own safety and evaluation tests.

Reward and evaluation

Rewards must be decomposed and logged rather than using only a single opaque score:

  • forward progress toward target: positive
  • heading/orientation accuracy: positive
  • upright/support stability: positive
  • commanded stop: positive when velocity settles
  • collision with obstacle: negative
  • unnecessary oscillation/reversal: negative
  • hazard exposure: strongly negative
  • incapacitation/death: terminal negative
  • action jerk and excessive motor effort: small negative

Report at least:

  • distance traveled and target progress
  • time upright
  • stop accuracy and turn error
  • oscillation frequency/amplitude
  • hazard avoidance rate
  • incapacitation/death rate
  • episode duration and simulation time
  • CPU allocations and frame time in the game

Compare against a non-neural hand-authored baseline and an untrained adapter baseline. Keep held-out seeds and scenarios so training cannot simply memorize one layout.

Runtime and asset requirements

  • Add an adapter asset format with a schema version, connectome payload hash, dimensions, parameter ranges, and training metadata.
  • Refuse to load an adapter trained for a different connectome hash or incompatible runtime schema.
  • Keep training artifacts out of the mod package unless they are the selected validated runtime asset.
  • Include a clear way to revert to the baseline adapter without deleting the person or game state.
  • Bound all queues, episode buffers, replay state, and telemetry.
  • Avoid per-tick allocations in the game path and measure full-graph startup and steady-state cost.

Testing requirements

  • Unit tests for observation normalization, encoder parameters, decoder math, clamping, smoothing, and reward terms.
  • Determinism tests for the brain plus adapter.
  • Serialization round-trip and corrupt/incompatible asset rejection tests.
  • Property tests for finite outputs, bounded actions, and no NaN/Infinity propagation.
  • Headless curriculum regression tests with fixed seeds.
  • People Playground manual tests for one and multiple controlled people.
  • Performance measurements for startup, per-tick CPU time, allocations, and memory.
  • Verify that all game-facing source files compile against the exact People Playground assemblies used by the installed game.

Suggested implementation order

  1. Complete or land the modular boundaries from issue Refactor runtime into modular core and body adapters #4.
  2. Extract PersonObservation and MotorCommand into the shared core contract.
  3. Add a deterministic headless environment and replay format.
  4. Add decoder-only training and compare it with the current heuristic adapter.
  5. Add the locomotion curriculum and evaluation reports.
  6. Add encoder tuning for the minimum sensory set.
  7. Add hazard/effect curriculum and held-out scenarios.
  8. Integrate the selected adapter asset into the game mod.
  9. Profile and test in People Playground.
  10. Add limb/grip interaction only after locomotion is stable.

Acceptance criteria

  • A clean checkout can run the headless training/evaluation commands with documented prerequisites.
  • Training is reproducible from a seed and produces a versioned adapter asset.
  • The runtime rejects malformed, incompatible, or non-finite adapter assets.
  • The trained adapter beats the current heuristic baseline on held-out locomotion scenarios.
  • In People Playground, one spawned Person Connectome human can stand, walk toward a target, stop, and turn without the observed persistent wiggle.
  • Hazard tests produce consistent avoidance/stop behavior without crashes.
  • The mod remains responsive within a documented CPU/memory budget.
  • Results, limitations, and exact training configuration are committed with the adapter asset.

References

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions