You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
This repository was archived by the owner on Sep 29, 2026. It is now read-only.
Train the body adapter around the fixed connectome
Summary
The current PersonConnectome controller loads a MaleCNS-derived connectome, but the fly neural circuit is being driven through heuristic human-body mappings. In-game testing shows neural activity and hazard responses, while ordinary locomotion mostly produces back-and-forth wiggling. We need a reproducible training pipeline for the sensory encoder and motor readout so the fixed connectome can control a People Playground human more effectively.
This issue is about training the adapter around the connectome, not blindly changing millions of biological synapses.
Dependency on the modular refactor
This issue depends on issue #4. The modular refactor remains required and is not replaced by this training work. Land the Core/Adapters/Presentation/Composition boundaries first, then add training through those interfaces.
Training must not introduce a second brain/runtime implementation, duplicate the connectome loader, or put optimizer logic into the People Playground entry point. The trainer should modify versioned encoder/readout parameters between episodes while the fixed connectome and game-facing adapter remain explicit and testable.
Before this issue is considered complete, #4 must provide stable contracts for PersonObservation, SensoryFrame, NeuralMotorFeatures, and MotorCommand, with one authoritative RuntimeBrain shared by runtime and tests.
Goals
Make stable locomotion the first trained behavior: stand, walk forward, stop, and turn.
Preserve the MaleCNS graph and its provenance by default.
Learn or tune the mapping from People Playground observations to sensory input populations.
Learn or tune the mapping from selected neural outputs to human-body actions.
Support hazards and reactive behavior: avoid lava/fire/electricity, recover from disturbances, and stop when incapacitated.
Provide a deterministic offline training/evaluation path before using the game runtime.
Export a versioned, validated adapter asset that the mod can load at runtime.
Non-goals
Do not retrain every connectome synapse as the first approach.
Do not claim biological equivalence between a fly brain and a human body.
Do not add an external Python, network, or cloud dependency to the shipped mod.
Do not make weapon use the first milestone. Gripping, aiming, and trigger control can follow once locomotion is reliable.
Proposed architecture
PersonConnectome.Core
MaleCNS asset/loader
LIF simulation
deterministic stepping
state/checkpoint support
PersonConnectome.Training
episode/environment interface
sensory encoder parameters
motor readout parameters
reward functions
mutation/optimizer strategy
checkpoints and evaluation reports
PersonConnectome.PeoplePlayground
human observation adapter
liquid/status/damage sensing
action application
game-time telemetry
PersonConnectome.Runtime
loads the trained adapter asset
validates version, graph hash, ranges, and dimensions
runs the fixed brain plus adapter in-game
The connectome should be immutable during normal training. Training parameters should be explicit and separate, for example:
input population IDs and gains
input normalization/clamping
temporal filters and decay
motor population IDs
left/right/forward/stop readout weights
limb/grip gains
action smoothing and dead zones
optional recurrent adapter state, with bounded size
Observation and action model
Record a fixed-rate PersonObservation frame containing only data available from the supported game API:
position, velocity, facing, angular velocity, and body tilt
head/limb positions and joint angles where available
liquid identity, wetness, exposure duration, and estimated severity
heat/cold, fire/lava, electricity, oxygen, blood, infection, poison, anesthesia, and injury state
target direction and optional curriculum objective
Decode a bounded MotorCommand such as:
forward/backward locomotion
turn left/right
stop/brake
per-limb target or motor effort
grip open/close
optional recovery/brace action
Every action must be clamped and rate-limited before reaching game APIs.
Training stages
Stage 0: deterministic harness
Add a headless environment with a simplified 2D body model and the same observation/action contracts.
Add fixed seeds, replayable episodes, checkpoints, and golden evaluation traces.
Verify that identical asset, adapter, seed, and episode produce identical outputs.
Stage 1: locomotion curriculum
Train in increasing difficulty:
remain upright and reduce unwanted oscillation
hold a heading
walk toward a target direction
stop on command
turn toward a target
walk around simple obstacles
recover from small pushes and uneven contact
Keep the connectome fixed. Start by optimizing only decoder weights and gains; add encoder parameters only when the decoder-only baseline cannot solve a stage.
Stage 2: sensory grounding
Compare learned inputs against hand-authored baselines. Train the smallest useful input set first:
forward/left/right target direction
support/upright error
velocity error
touch/collision
threat/heat/fire/lava
Add light, sound, chemicals, and other effects only after locomotion has a stable baseline. Log which input populations and output populations actually contribute to the behavior.
Stage 3: reactive survival
Add rewards and penalties for hazard avoidance, safe stopping, recovery, and survival. Test water, acid, poison, anesthesia, electricity, fire/lava, damage, blood loss, infection, and unconsciousness independently and in combinations.
Stage 4: limbs and interaction
After locomotion passes its acceptance thresholds, train reaching, gripping, releasing, and simple object interaction. Treat weapon/trigger behavior as a separate opt-in experiment with its own safety and evaluation tests.
Reward and evaluation
Rewards must be decomposed and logged rather than using only a single opaque score:
forward progress toward target: positive
heading/orientation accuracy: positive
upright/support stability: positive
commanded stop: positive when velocity settles
collision with obstacle: negative
unnecessary oscillation/reversal: negative
hazard exposure: strongly negative
incapacitation/death: terminal negative
action jerk and excessive motor effort: small negative
Report at least:
distance traveled and target progress
time upright
stop accuracy and turn error
oscillation frequency/amplitude
hazard avoidance rate
incapacitation/death rate
episode duration and simulation time
CPU allocations and frame time in the game
Compare against a non-neural hand-authored baseline and an untrained adapter baseline. Keep held-out seeds and scenarios so training cannot simply memorize one layout.
Runtime and asset requirements
Add an adapter asset format with a schema version, connectome payload hash, dimensions, parameter ranges, and training metadata.
Refuse to load an adapter trained for a different connectome hash or incompatible runtime schema.
Keep training artifacts out of the mod package unless they are the selected validated runtime asset.
Include a clear way to revert to the baseline adapter without deleting the person or game state.
Bound all queues, episode buffers, replay state, and telemetry.
Avoid per-tick allocations in the game path and measure full-graph startup and steady-state cost.
Testing requirements
Unit tests for observation normalization, encoder parameters, decoder math, clamping, smoothing, and reward terms.
Determinism tests for the brain plus adapter.
Serialization round-trip and corrupt/incompatible asset rejection tests.
Property tests for finite outputs, bounded actions, and no NaN/Infinity propagation.
Headless curriculum regression tests with fixed seeds.
People Playground manual tests for one and multiple controlled people.
Performance measurements for startup, per-tick CPU time, allocations, and memory.
Verify that all game-facing source files compile against the exact People Playground assemblies used by the installed game.
Train the body adapter around the fixed connectome
Summary
The current PersonConnectome controller loads a MaleCNS-derived connectome, but the fly neural circuit is being driven through heuristic human-body mappings. In-game testing shows neural activity and hazard responses, while ordinary locomotion mostly produces back-and-forth wiggling. We need a reproducible training pipeline for the sensory encoder and motor readout so the fixed connectome can control a People Playground human more effectively.
This issue is about training the adapter around the connectome, not blindly changing millions of biological synapses.
Dependency on the modular refactor
This issue depends on issue #4. The modular refactor remains required and is not replaced by this training work. Land the Core/Adapters/Presentation/Composition boundaries first, then add training through those interfaces.
Training must not introduce a second brain/runtime implementation, duplicate the connectome loader, or put optimizer logic into the People Playground entry point. The trainer should modify versioned encoder/readout parameters between episodes while the fixed connectome and game-facing adapter remain explicit and testable.
Before this issue is considered complete, #4 must provide stable contracts for PersonObservation, SensoryFrame, NeuralMotorFeatures, and MotorCommand, with one authoritative RuntimeBrain shared by runtime and tests.
Goals
Non-goals
Proposed architecture
The connectome should be immutable during normal training. Training parameters should be explicit and separate, for example:
Observation and action model
Record a fixed-rate
PersonObservationframe containing only data available from the supported game API:Decode a bounded
MotorCommandsuch as:Every action must be clamped and rate-limited before reaching game APIs.
Training stages
Stage 0: deterministic harness
Stage 1: locomotion curriculum
Train in increasing difficulty:
Keep the connectome fixed. Start by optimizing only decoder weights and gains; add encoder parameters only when the decoder-only baseline cannot solve a stage.
Stage 2: sensory grounding
Compare learned inputs against hand-authored baselines. Train the smallest useful input set first:
Add light, sound, chemicals, and other effects only after locomotion has a stable baseline. Log which input populations and output populations actually contribute to the behavior.
Stage 3: reactive survival
Add rewards and penalties for hazard avoidance, safe stopping, recovery, and survival. Test water, acid, poison, anesthesia, electricity, fire/lava, damage, blood loss, infection, and unconsciousness independently and in combinations.
Stage 4: limbs and interaction
After locomotion passes its acceptance thresholds, train reaching, gripping, releasing, and simple object interaction. Treat weapon/trigger behavior as a separate opt-in experiment with its own safety and evaluation tests.
Reward and evaluation
Rewards must be decomposed and logged rather than using only a single opaque score:
Report at least:
Compare against a non-neural hand-authored baseline and an untrained adapter baseline. Keep held-out seeds and scenarios so training cannot simply memorize one layout.
Runtime and asset requirements
Testing requirements
Suggested implementation order
PersonObservationandMotorCommandinto the shared core contract.Acceptance criteria
References