Skip to content

Latest commit

 

History

675 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

JIN Core Engine

Python FastAPI WebSocket OpenAI Compatible MCP Compatible Tests

JIN Core Engine is an experimental cognitive runtime for OpenAI-compatible models with visible memory, session continuity, and model-driven actions.

Built for long-running interaction, JIN keeps the context shaping each response inspectable while exposing memory, reasoning, runtime actions, persistent files, session restore, MCP skills, telemetry, and the Live Avatar without turning the main chat into a control panel.

Interface

JIN Core Engine runtime workspace

The JIN workspace combines the chat stream, draggable/collapsible runtime panels, runtime actions, persistent files, and the Live Avatar.

First Run / Install

The default Windows setup is one-click. You do not need to install Python, LM Studio, llama.cpp, or a model manually.

  1. Download or clone the repository and extract it to a normal writable folder.
  2. Double-click:
JIN_LAUNCHER.bat
  1. On the first run, the launcher first checks http://127.0.0.1:1234 for an already running LM Studio server:
    • If LM Studio is available and exposes models, JIN immediately creates config.py for that endpoint, skips the bundled llama.cpp and Gemma downloads, and opens the normal launcher dashboard with the detected models ready to choose. Select the Brain model and press Enter.
    • If LM Studio is not available, JIN falls back to the fully self-contained setup: it prepares a private Python 3.12 runtime, downloads and verifies the bundled llama.cpp CUDA runtime and default Gemma 4 E4B Instruct Q4_K_M model, then creates config.py, starts the local Brain/backend, and opens http://127.0.0.1:8000.

The launcher shows setup progress directly in its window when the embedded fallback is needed. Internet access is required only for components that are not already available locally. The bundled embedded-Brain path currently targets Windows x64 and uses the CUDA 12.4 llama.cpp build. Other platforms or external OpenAI-compatible model servers can use the manual/custom setup described below.

After the first successful run, start JIN with the same JIN_LAUNCHER.bat. If config.py already exists when the launcher starts, the first-run detection/bootstrap is skipped entirely: the normal dashboard appears immediately while the configured runtime is brought online.

config.py remains the persistent startup-mode switch. Delete it only when you intentionally want JIN to run first-start detection again: it will reuse LM Studio at 127.0.0.1:1234 when available, otherwise it will start the embedded bootstrap.

Live Avatar

Live Avatar visualizes JIN's runtime state in real time.

Inner orbits react to live FRAME/runtime-memory changes, while outer signal rings track Delayed Memory, L-T facts, Active Memory, and persistent files.

The non-rotating scaffold rings and breathing rays also mirror context pressure: they use the same green-to-warm progress color as the context meter, while ray peak opacity rises from roughly 0.10 toward 0.70 as the window fills. Rays fade fully out and back over a 30-second cycle.

The avatar is interactive: reasoning references light up matching runtime signals, memory-row hover zooms/highlights the corresponding signal, and larger L-T stores fan out across additional outer rings. The center toggle fades all scaffold/runtime/memory/file rings, then removes those hidden layers from painting/animation after the fade; the central light remains.

During reasoning, the avatar shifts into a dedicated motion state. Runtime actions can change its color, reaction, size, position, and speed, giving the model a small visual language beyond text.

Live Avatar memory rings

Memory Architecture

The memory panel has five views — FRAME, ACTIVE, DELAYED, L-T, and FILES — plus a LOGS archive view that projects saved sessions.

Memory panel

FRAME

FRAME is the live runtime-memory snapshot. It keeps the current topic, request/task state, decisions, feedback, and unresolved points needed by upcoming turns. Accepted updates are versioned as snapshots so the UI can step through diffs and inspect what changed. The latest FRAME value can also be edited directly from its memory tooltip; historical frames remain read-only. FRAME values follow the detected language of the current user message while structural keys remain English snake_case.

Active Memory

Active Memory keeps unfinished intentions and pending commitments separate from the general conversation state. Conditions and unresolved contracts remain active across turns until they are fulfilled, cancelled, paused, or explicitly resolved. Relevant active records are projected back into Brain context without changing their canonical storage order. Conditions can be edited directly from the inspector while IDs, keys, custom fields, and status metadata remain structurally owned by the runtime.

Delayed Memory

Delayed Memory stores larger structured context that should be available without living in every prompt. Reports can link L-T facts and persistent files, can be loaded/unloaded by runtime actions, pinned from the UI, or surfaced from matching user-text tags. Panel rows expose a compact hover preview with summary, tags, IDs, linked facts, creation time, and a bounded body preview; unpinning is also represented in the shared memory logger flow.

Long-Term Facts

L-T is the UI view of durable facts: stable user/project facts, preferences, constraints, decisions, and environment details that should survive sessions. An internal candidate buffer feeds idle extraction and merge. Facts absorbed into Delayed reports stay hidden from the default active view but can be revealed with the count toggle; report-linked fact IDs open the owning report. Explicit fact values are editable, and fact mentions refresh recall so recently used facts stay fully expanded in Brain context while older facts fall back to compact sentence previews.

Files

FILES exposes the persistent uploaded-file library. Stored files keep stable IDs and can be attached/detached across turns or linked from Delayed Memory. Files attached to the next message also appear as compact composer chips: click to preview, hold to detach from context without deleting the stored file.

Logs

LOGS is the disk-backed archive browser for restorable sessions. Session titles come from the protected session_title field inside FRAME and update in the list when a newly committed FRAME changes the title; older archives without a title fall back to their session ID. Hovering a row loads a bounded preview of the newest USER/JIN turns, while a normal click opens that archived session through the existing restore flow in a new tab. Holding a row for 1.5 seconds uses the shared fade/delete interaction to remove that saved session from disk; an empty date directory is removed only when nothing else remains inside it. Anonymous and greeting-only/technical sessions are hidden from the archive.

Core Capabilities

  • Inspectable Memory: Keeps FRAME/live state, long-term facts, delayed reports, active commitments, persistent files, and the LOGS session archive as distinct systems.
  • Session Continuity: Supports in-process soft WebSocket resume plus disk-owned reload/new-tab/bootstrap continuity and explicit archived-session restore from persisted logs.
  • Persistent Files: Stores uploaded text, images, PDFs, and other files under stable ids; the same stored files can be attached to or detached from context across turns.
  • Runtime Telemetry: Shows model status, token usage, live context pressure, memory updates, action state, and runtime logs; at 50%+ previous-answer context usage the Brain also receives a live <CONCERNS> warning, with an explicit cleanup reminder when tool results are present. Empty concerns are omitted. The status modal can switch the configured LM Studio model for an available runtime role.
  • Reasoning highlighting: Displays provider/model reasoning separately from the final answer when the backend exposes a reasoning stream. Maps direct references back to runtime rules, memory records, restored context, and linked runtime objects.

Reasoning citation highlighting

Runtime Actions

JIN can request an action while answering. The runtime validates and executes it, then returns any required result to the model before the workflow continues. Concrete contracts carry their own readable schema; a failed action is returned as a human-readable tool result with the reason, supplied payload when relevant, and the correct schema, followed by an explicit continuation instruction so the Brain does not treat the failed mutation as completed. Runtime execution preserves the model's emitted source order: each action is prepared and run before the next one; contract runtime_order only affects how action instructions are listed to the model.

Current contract families include:

  • SAVE_ACTIVE_MEMORY for both create/update, plus paired multi-ID DELETE_ACTIVE_MEMORY;
  • SAVE_DELAYED_MEMORY and paired multi-ID LOAD_DELAYED_MEMORY; loaded reports are removable tool results, while only user-pinned reports enter the dedicated loaded-memory block;
  • LIST_ALL_USER_SHARED_FILES, plus skill-gated ATTACH_FILE_CONTENT and paired multi-ID ATTACH_FILES_BY_ID (internally ATTACH_FILE_BY_ID per file);
  • LOAD_SKILLS_CONTEXT and UNLOAD_SKILLS_CONTEXT accept comma-separated skill lists; the internal actions remain singular;
  • POSTING_BOARD after the posting_board skill is loaded;
  • CALL_MCP after any skill containing a valid <MCP_SERVER>...</MCP_SERVER> declaration is loaded;
  • JIN_COLOR, JIN_REACTION, JIN_SIZE, JIN_POSITION, and JIN_SPEED.
  • WEB_SEARCH, DEEP_WEB_SEARCH, and local CHAT_LOG_SEARCH when their capability gates allow them;
  • CLEAN_TOOL_RESULTS, UPDATE_LT_FACTS, and model-facing <RECALL_FACTS_CONTEXT> F1, F2 </RECALL_FACTS_CONTEXT>;

Concrete schemas in contracts/*.json are authoritative.

Architecture

JIN Core Engine architecture

Runtime Flow

The WebSocket layer resolves a session-owned RuntimeContext for the browser client. A soft reconnect can reattach to the same live runtime/transport instead of creating a new foreground state container; explicit page departure retires it, while an unexplained disconnect has a 600-second reconnect grace. Every accepted user message is then handled by AgentRuntime.

A normal turn follows this path:

  1. The user sends a message with optional persistent attachments.
  2. AgentRuntime passes the request directly to the Brain.
  3. The Brain streams reasoning and visible answer content through separate runtime channels.
  4. Stream validation guards repetition and malformed generation while private runtime-action markers are extracted.
  5. Runtime Actions execute in model-emitted source order, can mutate state or return trusted results, and actions that need another model step continue inside the same user sequence.
  6. After the visible turn completes, the logical Service route performs background FRAME integration; if no dedicated Service endpoint is configured, this route reuses the Brain client.
  7. A later user turn waits for any pending FRAME update, then receives current Active Memory, <FRAME_MEMORY_N> followed by up to five recent USER/JIN pairs, loaded Delayed Memory, L-T facts, files/skills, action history, context-usage/concern signals, and trusted tool results.

The model path is intentionally direct:

user -> brain

Planning decisions, runtime actions, and follow-up decisions all happen inside the Brain/runtime loop.

Model Roles

JIN talks to models through an OpenAI-compatible API.

The runtime separates model work into roles:

  • Brain: visible reasoning, responses, and runtime decisions;
  • Service: background memory updates and supporting work.

JIN is model-agnostic at the API boundary. Brain is the only foreground response route. Service is background-only. SERVICE_API_BASE is optional: when it is empty, the Service client aliases the Brain client, so one physical model can handle both logical roles without changing foreground routing. Set SERVICE_API_BASE only when a dedicated background Service node exists.

On Windows, the LM Studio launcher can fill unset/default Brain model settings from a loaded Gemma-family model and can separately initialize a dedicated Service endpoint when one is configured. Explicit provider URLs and model ids remain unchanged.

Runtime Storage

Reload/bootstrap authority is disk-owned. Browser cognitive state is a page-local projection: the live jin.liveRuntimeMemory.v2 record is cleared whenever the page module starts. A soft WebSocket reconnect can reuse the surviving server RuntimeContext; after a backend/page restart JIN rebuilds continuity from disk.

Persistent state is stored through:

  • logs/YYYY-MM-DD/<session>/ for USER/JIN dialogue, reasoning, runtime events, server checkpoint/tool-result events, and saved frames/ snapshots used by normal bootstrap and archived restore;
  • logs/.continuation-cleared.json for the USER-count barrier created by Session CLEAR;
  • memory/active/*.json for Active Memory;
  • memory/delayed/*.json for Delayed Memory reports;
  • memory/facts/long_term_facts.json plus pending_facts.json for durable L-T and its candidate queue;
  • assets/files/ plus its local index for persistent uploaded files.

UI preferences may use browser storage. Model and search traffic goes to the endpoints and providers configured for the runtime.

Assets and Skills

Reusable material lives under assets/:

assets/
|-- skills/       # Instructions and optional local Python tools
|-- files/        # Persistent uploaded-file library
|-- prompts/      # Reusable prompt lists
|-- templates/    # Prompt templates
|-- wildcards/    # Text values used by templates and generators
`-- outputs/      # Generated files

JIN can inspect <SKILLS_LIST>, load required skills with one comma-separated <LOAD_SKILLS_CONTEXT> ... </LOAD_SKILLS_CONTEXT> block, run their allowed actions, and unload one or more with <UNLOAD_SKILLS_CONTEXT> ... </UNLOAD_SKILLS_CONTEXT>. Loaded skill bodies are projected through the normal tool-results context, while <SKILLS_LIST> remains the compact availability/loaded-state inventory. Python skills execute from .py files inside the selected skill directory with bounded execution and output limits. Persistent uploaded files are stored separately under assets/files/ and keep stable ids across turns.

MCP skills

JIN is an MCP client for tool servers. A skill can declare one MCP server in its JIN_SKILL.md with a machine-readable <MCP_SERVER>...</MCP_SERVER> JSON block. Loading that skill opens/discovers the server, appends the live tools/list catalog to the in-memory skill context, and enables the single generic <CALL_MCP>...</CALL_MCP> runtime action. Tool-specific names and argument schemas stay in the skill/server; adding another MCP integration does not require another Python runtime action.

Supported transports are stdio, Streamable HTTP, and SSE (http / streamable-http normalize to Streamable HTTP). Stdio connections remain alive across automatic JIN follow-ups and are closed when the skill/runtime is unloaded; an optional positive read_timeout_seconds applies to all supported transports. MCP image results are stored in the normal JIN file store and injected as image attachments into the next Brain follow-up, so visual tools can return screenshots/renders without embedding base64 into <TOOL_RESULT>. Generic MCP bubbles open the structured request/result trace, while get_viewport_screenshot reuses the normal attachment preview. See docs/MCP_SKILLS.md for the skill contract.

Project Layout

.
|-- app.py                     # FastAPI app, routes, and lifespan
|-- websocket/                 # WebSocket routing, messages, and UI logging
|-- contracts/                 # Action markers, rules, guards, and follow-ups
|-- agent/                     # Direct Brain runtime, state, and Brain node
|-- clients/                   # OpenAI-compatible client builders
|-- runtime/                   # Context, memory, streams, telemetry, registry
|-- memory/                    # Runtime-created Active/Delayed/L-T stores (gitignored data)
|-- assets/                    # Skills, persistent files, prompts, and generators
|-- rules/                     # Brain and runtime rule blocks
|-- utils/                     # Actions, assets, validation, and storage helpers
|-- ui/                        # Browser interface and README images
|-- tests/                     # Unit, action, and model-integration tests
|-- config.example.py          # Configuration template
|-- config_loader.py           # Local configuration loader
|-- app_settings.py            # Typed settings wrapper
|-- JIN_LAUNCHER.bat           # Windows one-click launcher
|-- jl.ps1                     # Windows bootstrap, runtime, and launcher UI
|-- Dockerfile                 # Container image
|-- compose.yml                # Docker Compose local runtime
|-- requirements.txt           # Python dependencies
|-- package.json               # Test and probe commands
|-- docs/                      # Current architecture, state, and durable decisions
`-- LIVE_AVATAR.md             # Avatar visual-state contract

Advanced Setup

The one-click Windows launcher above is the recommended path. The options below are for custom providers, non-default environments, manual startup, or containers.

Custom / Manual Requirements

  • Python 3.10+ when starting JIN manually
  • One or more OpenAI-compatible model servers when not using the embedded Windows Brain
  • Node.js 20+ only for local tests and behavior probes
  • A Serper API key only when built-in web search is enabled

An external model server must expose:

/v1/chat/completions
/v1/models

For LM Studio, JIN also probes the provider-native /api/v1/models metadata endpoint and falls back to legacy /api/v0/models when needed, allowing the runtime to read the context length of the model that is actually loaded.

Using an Existing OpenAI-Compatible Brain

If you want to use LM Studio or another compatible server instead of the bundled local Gemma runtime, create config.py from config.example.py before launching JIN and set BRAIN_API_BASE to that server. An explicit Brain URL makes the Windows launcher preserve that configuration and skip the embedded llama.cpp/model bootstrap. BRAIN_MODEL_UID may be left empty so the launcher can discover the endpoint's model catalog.

A single model is enough by default: with SERVICE_API_BASE left empty, background Service work reuses Brain. Configure a second endpoint only if you want a dedicated Service model.

Then run:

JIN_LAUNCHER.bat

Manual Start

git clone https://github.com/makeitdouble/jin_core.git
cd jin_core
python -m venv .venv

Windows PowerShell:

.\.venv\Scripts\Activate.ps1
Copy-Item config.example.py config.py

Linux/macOS:

source .venv/bin/activate
cp config.example.py config.py

Before starting, edit config.py if your provider URLs or model ids differ from the template values.

Install and run:

pip install -r requirements.txt
python app.py

Then open:

http://127.0.0.1:8000

Docker / Docker Compose

Docker uses the same config.py and .env contract as the native launcher. The Compose setup bind-mounts config.py, memory, logs, persistent files, and generated outputs so recreating the container does not reset JIN.

Make sure config.py and .env exist, then run:

docker compose up --build

Open:

http://127.0.0.1:8000

By default Compose points BRAIN_API_BASE at http://host.docker.internal:1234, because 127.0.0.1 inside the container means the container itself. To use another Brain endpoint, set BRAIN_API_BASE in the root .env before starting Compose. A configured dedicated SERVICE_API_BASE continues to come from config.py unless you override it through the environment.

If you use JIN's linked-project-folder feature, remember that the backend can only see paths mounted into the container; add an extra bind mount for any external project directory you want JIN to read.

Stop the container with:

docker compose down

Configuration

The Windows one-click launcher creates config.py automatically after a successful first run. For manual or external-provider setups, copy config.example.py to config.py yourself and set the provider URLs and model IDs. config.py is ignored by Git.

Option Purpose
ENABLE_RUNTIME_LOGS Enable local runtime/chat logs.
BRAIN_API_BASE, BRAIN_MODEL_UID, BRAIN_TEMPERATURE Configure the required foreground Brain runtime.
BRAIN_MAX_FOLLOWUPS Limit executable internal action/follow-up ticks per user turn. Positive values cap the workflow and then allow one final non-executable response tick; 0 means unlimited follow-ups.
SERVICE_API_BASE, SERVICE_MODEL_UID, SERVICE_TEMPERATURE Optionally configure a dedicated background Service runtime. Leave SERVICE_API_BASE empty to reuse Brain.
LT_IDLE_SECONDS Set the L-T background consolidation idle delay. L-T memory itself is always enabled.
SEARCH_PROVIDER, SEARCH_MAX_RESULTS Configure the built-in web-search provider and result count.
DEEP_WEB_SEARCH_MAX_QUERIES_PER_WORKER, DEEP_WEB_SEARCH_MAX_WORKER_CALLS Bound worker fan-out/call count for Deep Web Search.

User-facing config values can also be supplied through environment variables. Plain names and JIN_-prefixed names are supported; plain names take priority.

For optional credentials used by the Windows one-click launcher, copy .env.example to .env in the repository root and replace only the placeholders you need. JIN_LAUNCHER.bat delegates to jl.ps1, which loads that file into the JIN process before configuration is resolved. Variables already present in the process environment are not overwritten. The local .env is ignored by Git; .env.example contains names and placeholders only and is safe to commit.

SEARCH_SERPER_API_KEY=your-serper-api-key
GETPOSTINGBOARD_API_KEY=your-getpostingboard-api-key

When starting JIN manually with python app.py, export the same variables in the shell first; automatic .env loading belongs to the Windows launcher.

Secrets are environment-only and are intentionally not stored in config.py:

  • SEARCH_SERPER_API_KEY or JIN_SEARCH_SERPER_API_KEY
  • GETPOSTINGBOARD_API_KEY or JIN_GETPOSTINGBOARD_API_KEY

Tests

Run the local suite:

npm test

You can also run the same suite directly with Python:

python -m tests.run_unittest

Run optional behavior probes:

npm run probe ascii
npm run probe movie
npm run probe word
npm run probe marker
npm run probe save
npm run probe delayed

GitHub Actions runs the same suite. Browser client tests stay separate because they require Playwright and an Edge browser channel; run them explicitly with npm run browser_tests in an environment that provides those dependencies. Model-dependent probes remain local unless CI is connected to a compatible runtime.

About

An experimental cognitive runtime with visible memory.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages