Skip to content

About

A personal AI agent with personality and memory managed by an external shell. The agent lives 7 days; the diary emerges from that life.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

English · 中文 · 日本語

emergent-diary-agent

A personal AI agent with personality and memory managed by an external shell. The agent lives 7 days; the diary emerges from that life.

flowchart TB
    Boot["Day 0: Bootstrap<br>~200-char inputs → Personality, World, Voice Bible"]

    subgraph Daily ["Daily Loop · Day 1 → Day 7"]
      direction LR
      W["1. World Engine"]
      E["2. Day Experience"]
      R["3. Evening Reflection"]
      C["4. Composer + Editor"]
      N["5. Night Consolidation"]
      W --> E --> R --> C --> N
    end

    Stores[("3 persistent JSON stores<br>Character · World · Memory<br>(6-layer cumulative)")]

    Boot --> Daily
    N -. writes .-> Stores
    Stores -. reads .-> W
Loading

▶ Open the live showcase  ·  📖 Read the 7-day diary  ·  💻 Run with your own character

⚠️ Output language is Japanese. The shipped run produces Japanese diaries, and the showcase HTML is in Japanese. The architecture itself is general-purpose, but the SYSTEM prompts are currently Japanese-locked. See Language support for what it would take to switch.


What this project is

This is a Python AI character system. Feed it a ~200-character description of a person and a ~200-character description of an initial situation, and it makes a fictional character with a defined personality live through 7 fictional days and write seven diary entries about them.

Every day, the character goes through five stages:

  1. World Engine independently generates 2-4 scenes for the day; the world moves on its own, regardless of the character's wishes;
  2. Day Experience runs the scenes sequentially through the character's personality filter, producing emotionally-coloured subjective memories;
  3. Evening Reflection has the character ruminate at dusk: what stuck, what got skipped, what tone today's diary should carry;
  4. Composer + Editor transcribes the now-personalized inner state into a diary entry, with a lint-triggered editor that rewrites when AI-smell residue is detected;
  5. Night Consolidation integrates today's experience into a 6-layer long-term memory, so tomorrow starts from a continuous self.

The seven diary entries are what falls out of the back of stage 4. They are not what the model is being asked to "write." That distinction looks subtle but it determines almost every architectural choice in this project.

If you just want to read the diary, open the live showcase — there is nothing to install. If you want to understand why every architectural choice is what it is, keep reading.

Three claims this project tests

1. Shell, not model. Personality continuity, memory governance, and world consistency are all implemented in code outside the LLM (the "shell"). The LLM remains a stateless next-token predictor; what makes the agent feel persistent is the shell. To test this, the same shell was run against both Claude Sonnet 4.6 (strong) and Claude Haiku 4.5 (weak). Structural properties (7-day memory continuity, belief deduplication, emotion diversity, anti-AI-smell closure) held on both. Prose ceiling differed; the shell did not.

2. Emergence, not generation. The diary is never the optimization target. The daily loop optimizes for living through the day with personality intact; the diary text is what falls out of that process. Stages 1-3 modify the agent's internal state; stage 4 (Composer) faithfully transcribes an already-personalized inner state. Its job is not to "write something interesting." This is how AI smell gets eliminated upstream rather than as post-hoc cleanup.

3. Continuity outlasts the model. Nine rounds of iteration produced a shell that survives model substitution. Same code, same character inputs — run through Haiku and Sonnet — produce two diaries with different prose quality but the same lived narrative arc, the same memory chain, the same world facts referenced consistently across days. The shell is the load-bearing component; the model is the prose ceiling.

How it works

Bootstrap (Day 0)

The system takes only two inputs from the human: a ~200-character description of the character and a ~200-character description of the initial situation. Nothing else is hardcoded. From these two, the system generates three things:

  • Character Profile: Big Five baseline, core values, cognitive biases, and an abstract expression-style description (a short paragraph describing how this character writes diaries; e.g., "short sentences, frequent 体言止め [noun-ending sentences], sensory-detail preference");
  • Initial World State: setting, NPCs, ongoing story threads;
  • Voice Bible: 2-3 paragraphs of actual sample diary prose in this character's voice, generated by combining the character profile with six pre-LLM-era Japanese diary style priors.

The expression-style description and the Voice Bible are layered: the former is an abstract spec (saying how this person writes), the latter is a concrete reference (showing what this person writes). The Composer reads both. The benefit of this layering is that the model can pick up the feel of the character's voice from concrete prose, rather than having to infer it from abstract bullets.

The five-stage loop (key invariants)

The five stages were introduced above. Here are the invariants that do the actual engineering work in the shell:

  • Daytime reads from a frozen snapshot. At the start of each day, profile / world / memory are copied to immutable snapshots; all five stages read from those snapshots. Only at stage 5 does the next-day state get written. This is what makes within-day consistency and across-day growth coexist without race conditions.
  • Emotion has momentum. new_valence = old × 0.7 + impact × 0.3. A bad day does not reset overnight.
  • The shell decides emotion labels, not the model. Letting weak models name emotions freely tends to collapse to a handful of labels (happy/sad/anxious). Here, valence/arousal coordinates are mapped to labels via a code-side dictionary. Diversity is preserved structurally.
  • Six memory layers, all diff-updated, with capacity caps. Episodic, semantic-self, beliefs, relationships, narrative state, and the diary archive — none are overwritten; capacity caps (concerns ≤ 3, hooks ≤ 5) force curation rather than infinite accumulation.
  • All three persistent stores are JSON files on disk. No database, no vector index. State is human-readable, diffable, version-controllable.

Language support

Because the shipped SYSTEM prompts hardcode the diary-writing language, only Japanese diary generation is currently supported. Switching to another language requires editing these eight files:

File Japanese-locked content
code/src/bootstrap.py SYSTEM prompt contains "すべて日本語で出力してください" and similar explicit instructions
code/src/world_engine.py Scene-generation prompt in Japanese
code/src/experience.py Experience-description prompt in Japanese
code/src/reflection.py Rumination prompt in Japanese
code/src/composer.py Diary-writing prompt in Japanese, plus Japanese-specific instructions like "drop the subject pronoun"
code/src/editor.py Contains a "## 日本語の自然さチェック" (Japanese naturalness check) section
code/src/consolidation.py Memory-integration prompt in Japanese
code/voice_style_priors.md Six pre-LLM-era Japanese diary style priors; would need to be replaced with target-language equivalents

The last one is the hardest. The six style priors are the shell's main anchor against the model drifting back into "assistant register." They have to be real, pre-LLM-era diary/essay/letter prose in the target language, exhibiting a noticeably different texture from generic AI prose. Curating equivalents for Chinese or English is non-trivial editorial work.

PRs welcome. If you fork this and get another language working end-to-end, please consider contributing a language-switch back to main.

Why this matters beyond diaries

The five-stage daily loop is, by itself, the minimal skeleton of a personal AI assistant.

Re-read each stage with the diary context stripped out:

  • World Engine = the channel through which a personal assistant receives signals from a world it doesn't control. Today: a generated text world. Tomorrow: real multimodal input such as the user's calendar, location, voice messages, photos.
  • Day Experience = the assistant reacting to those signals with personality intact. The same email read by an anxious helper and a laconic helper produces different memory traces.
  • Evening Reflection = the assistant ruminating during downtime: what to keep, what to forget, what to carry into tomorrow. This is what makes "long-term assistant" different from "long context window."
  • Composer + Editor = the assistant's outward expression layer. Here it is diary text; in a real assistant it could be a daily summary, a proactive nudge, a planning suggestion, or whatever output channel it has.
  • Night Consolidation = offline integration. Selective compression of today into long-term memory, so tomorrow starts from a continuous self.

Strip the diary, swap the world for multimodal input, add bidirectional action (the assistant decides whether to ping the user — the world responds), and what remains is the architectural skeleton of a personal AI assistant that grows alongside its user.

This project does not build that full assistant. It builds the shell that such an assistant would sit inside, and tests it through the simplest possible observation window: seven days, one persona, one prose register.

Built on the shoulders of

This project is not a clean-room invention; it stitches together several lines of recent agent and personality research, picking the load-bearing idea from each.

  • PIS — Personality Intelligence System (Wan et al., 2025). The thesis that personality can be implemented as a shell around the LLM, not as a property of the LLM — Big Five baseline → cognitive style → expression layer. This project implements the Big Five baseline portion; the Id/Ego/Superego middle layer is left as future work.
  • Hermes Agent (NousResearch). Frozen-snapshot session model and capacity-bounded memory curation. Both are directly used: each day reads from immutable snapshots, and active_concerns / unresolved_hooks have hard caps to force selection over accumulation.
  • OpenClaw / SOUL.md. File-as-memory architecture. All agent state is plain JSON on disk. No database, no vector store. State is human-readable, diffable, version-controllable.
  • Generative Agents (Park et al., 2023). Memory retrieval scored by recency × importance × relevance. Reflection as a periodic process distinct from observation.
  • PSM — Persona Shape Model (Anthropic, 2026). The diagnosis that AI smell is not a writing problem but a personality-collapse problem caused by RLHF flattening the persona space. This shaped the upstream-not-downstream approach to anti-AI-smell.
  • Claude Code's Auto-Dream pattern. Memory consolidation as an explicit, scheduled offline phase rather than inline with response generation. This project's Night Consolidation stage borrows that timing model.
  • LIWC. Big Five → linguistic markers mapping, used during Voice Bible generation to ground each character's prose style in measurable linguistic features rather than vibes.

What's inside

emergent-diary-agent/
├── code/                       Implementation: 14 Python modules
│   ├── src/                    bootstrap, world_engine, experience, reflection,
│   │                           composer, editor, consolidation, evaluator, ...
│   ├── voice_style_priors.md   Six pre-LLM-era Japanese diary style priors
│   ├── requirements.txt
│   └── .env.example
├── results/                    Two pre-computed full runs (no API key needed to inspect)
│   ├── sonnet_primary/         Claude Sonnet 4.6 — strong-model showcase
│   └── haiku_reference/        Claude Haiku 4.5 — weak-model A/B comparison
├── showcase/                   Static HTML viewer (data is pre-baked into JS)
├── tools/                      build_showcase.py + manifest, for regenerating showcase data
├── index.html                  Redirect → showcase/
└── 成果物(7日分の日記).md       The Sonnet run's 7-day diary, single readable file

Run it locally

The character is not hardcoded into the shell. Two short text descriptions are the only human-authored input; everything else (Big Five profile, world state, NPCs, Voice Bible) is generated from those two.

pip install -r code/requirements.txt
cp code/.env.example code/.env       # add your ANTHROPIC_API_KEY
# Edit code/src/main.py — only CHARACTER_INPUT and SITUATION_INPUT
# (each ~200 chars, in Japanese) need changing.
python code/src/main.py

Output goes to code/output/ and code/data/; this does not overwrite the pre-baked archives under results/.

Note:

  • The shipped run produces Japanese diaries only. Even if you supply a Chinese or English character description, the downstream prompts force Japanese output.
  • To get diaries in another language, see Language support above.

License

MIT — free for any use, including commercial.

About

A personal AI agent with personality and memory managed by an external shell. The agent lives 7 days; the diary emerges from that life.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages