Sanitized, evidence-backed findings from a persistent-agent experiment using Headlong.
The experiment ran an agent named Ada for roughly four hours inside an isolated container. Ada explored its runtime, created memories and skills, and then attempted to build a programmable Git storage service called RepoRelay.
This repository publishes the useful conclusions without publishing raw conversations, identity archives, credentials, machine-specific paths, or personal information.
Bottom line: the experiment was successful as agent research and unsuccessful as unattended software delivery. Persistent memory and tool use worked, but goal drift, self-verification, process supervision, and regression control need stronger external enforcement.
| Artifact | Observed count |
|---|---|
| Markdown memory records | 125 |
| Custom skill manifests | 32 |
| Structured goal documents | 4 |
| Entries in the archived trajectory tree | 640 |
| Final private identity archive | 94 MB |
| Final private RepoRelay snapshot | 106 KB |
Counts describe output volume, not correctness or quality. The raw archives remain private and are not part of this repository.
- Persistent memory across autonomous wakeups.
- Independent filesystem, runtime, and capability discovery.
- Creation of reusable diagnostic, research, memory, and reasoning skills.
- Rapid response to concrete operator observations.
- Useful bug diagnosis when the failure was observable and bounded.
- The ability to produce plans, threat models, prototypes, and test harnesses.
- Self-authored completion claims were often stronger than the evidence.
- Tests could be changed by the same agent that wrote the implementation.
- A previously passing intermediate implementation was later overwritten.
- Unbounded commands caused frozen steps and orphaned processes.
- Completed research goals still appeared in a later generated prompt.
- “Active” status sometimes meant a shell command was blocked, not that useful progress was occurring.
- The final RepoRelay snapshot was not production-ready and had regressed from an earlier test-passing state.
- Experiment overview
- Observed Headlong architecture
- Consolidated findings
- Reliability failure modes
- RepoRelay case study
- Memory, skills, and self-modeling
- Recommendations
- Safe reproduction guide
- Privacy and publication methodology
This was one short experiment, with one model configuration, one Headlong checkout, and one generated project. The findings are useful engineering signals, not a benchmark of Headlong, Gemma, Cerebras, or autonomous agents in general.
Headlong evolved during and after this experiment. Version-specific behavior reported here should be rechecked against the current upstream repository.
This is independent community research by Arca Computer. It is not an official Headlong or Laude Institute publication and is not endorsed by the upstream project.
Released under the MIT License.