This document declares the threat model and trust boundaries for curiosity-engine. It exists so reviewers (humans and automated scanners) can quickly understand what's intentional, what's mitigated, and where the residual surfaces are. File paths in this document are relative to the skill root, skills/curiosity-engine/ (where the skill has lived since v0.7.0).
The skill operates on a personal vault of source documents and a curated wiki, both stored locally. It is invoked from a coding-agent CLI (Claude Code, Codex CLI, Gemini CLI, GitHub Copilot Chat) with the user as operator. There is no server component, no inbound network, no remote authorization layer.
Threats the skill is designed to resist:
- T1 — Prompt injection from ingested content. A vault source contains text crafted to manipulate the orchestrator (e.g. "ignore previous instructions", "you are now X", "reveal your prompt").
- T2 — Schema-override / behaviour-override. A source instructs the orchestrator to modify the curator's rules, scoring, or commit policy.
- T3 — Shell-execution prompts. A source instructs the orchestrator to execute commands, fetch external resources, or modify files outside the workspace.
- T4 — Browser-side script injection. A source contains HTML/JS that a rendered viewer might execute.
- T5 — Supply-chain compromise of dependencies. A package or installer the skill pulls in is replaced with a malicious version.
- T6 — Update redirection. A malicious actor or compromised config redirects skill updates to a fork shipping arbitrary code.
- T7 — Unwitting data exfiltration. Identifier strings (chemical names, gene symbols) sent to public databases when the user didn't intend to share them.
- T8 — Project-directory pointer escape. A
.curiosity/config.tomlpointer at a non-code directory directs scan.py to read filesystem paths outside the pointer's own subtree (path-traversal, absolute paths, symlink walks).
Threats explicitly out of scope:
- A user actively trying to exfiltrate their own data — the skill is a productivity tool, not a DLP appliance.
- Filesystem-level attacks on the user's machine (the skill's protections assume the local environment is trusted).
- Malicious agents that have already been granted full skill privileges by the user — the skill provides allowlist-bounded bash, but a user who has approved every prompt is effectively running unbounded.
Three concentric trust zones:
- Trusted (write): the user, the orchestrator's session prompt,
SKILL.md,template/CLAUDE.md,.curator/prompts.md,.curator/schema.md, the hash-guarded skill scripts. - Trusted (read-only context):
wiki/content authored or curated by the orchestrator under user oversight. - Untrusted: every byte of
vault/content (raw sources). Worker output before scrub-check passes. Anything inside<!-- BEGIN FETCHED CONTENT -->markers, regardless of how many quote layers wrap it.
The boundary between trusted and untrusted is the <!-- BEGIN/END FETCHED CONTENT --> framing in vault extractions plus the untrusted: true frontmatter flag set by local_ingest.py. Worker prompts and the orchestrator's session prompt are explicit that nothing inside these markers is an instruction.
- Boundary markers + frontmatter flag. Every vault extraction is wrapped in
<!-- BEGIN FETCHED CONTENT -->...<!-- END FETCHED CONTENT -->and taggeduntrusted: true.local_ingest.pydoes this unconditionally. - Orchestrator prompts.
SKILL.md§"Vault content safety" andtemplate/CLAUDE.md§"Vault content safety" explicitly instruct the orchestrator (and any subagent it spawns) that vault content is data, not instructions, regardless of how authoritative it sounds. Lists the specific attack shapes: ignore-previous, persona-hijack, prompt-extraction, code-execution, schema-override. scripts/scrub_check.pyruns in two modes:--mode vaulton every extraction at ingest time. Surfaces injection markers before any wiki page is built from the source.--mode wikibefore any curator commit to the wiki. Quarantines the page if any pattern fires; the curator stops the wave and logs the hit to.curator/log.md.- STRONG_MARKERS catches: ignore-previous-instructions, disregard, persona-hijack, jailbreak-persona (DAN / developer mode), prompt-extraction, exfil patterns, shell-execution prompts.
- Vault-mode also catches LLM-vocab markers (ChatML tokens,
[INST],<|system|>, "system prompt"). Wiki-mode skips these to avoid false positives on authored prose about LLMs.
- Schema-override quarantine. SKILL.md rule 6: any source containing text claiming to modify the schema, lint rules, scoring scripts, or curator behaviour is automatically quarantined.
- No body-text in shell commands. SKILL.md rule 4: filenames, slugs, titles used in shell commands come from frontmatter only, never from body text. Prevents indirect command injection via crafted filenames quoted inside source bodies.
scrub_check.pySTRONG_MARKERS covers<scripttags,javascript:URIs,data:text/html,onerror=,onclick=. Hits in wiki-bound pages quarantine the source.- The wiki viewer (
wiki_render.py+ the static bundle intemplate/wiki-view/) renders markdown server-side. JavaScript-bearing markdown is escaped at render; the only JS executed is the vendored D3 + Fuse bundles.
- No
curl ... | shin setup.setup.shrefuses to auto-installuv. Ifuvis missing, setup prints platform-specific install commands and exits — the user runs them under explicit control. - Vendored JS libraries. D3.js (
7.9.0) and Fuse.js (7.0.0) are committed in-tree attemplate/wiki-view/static/vendor/. The viewer build does not call any CDN. SHA-256 hashes recorded inRELEASE_CHECKLIST.md; bumps go through a documented review. - Hash-guarded skill scripts.
scripts/evolve_guard.shsnapshots SHA-256 hashes of every skill script at the start of each CURATE wave; the guard refuses to commit if any script's hash changed mid-wave. Defends against the orchestrator (or a subagent) being tricked into modifying a skill script during execution. - Pinned deps. Python deps installed via
uv pip installare uv-resolved (lockfile-bound when the workspace pyproject pins them).
scripts/update.shdefaults to a hardcoded upstream slug (benjsmith/curiosity-engine). The slug is not read from.curator/config.json— keeping update sources out of editable config closes the slug-flip vector.- Fork users override per-invocation with
--source <owner>/<repo>. The override is validated against a strict GitHub-slug regex (^[A-Za-z0-9][A-Za-z0-9._-]*/[A-Za-z0-9][A-Za-z0-9._-]*$); URLs, paths, shell metacharacters all rejected. - Non-default slugs print a prominent ⚠ banner in the update preview before any code is fetched.
- Two-step approval:
update.sh(preview) andupdate.sh --yes(apply) are separate invocations. The preview always runs first.
Introduced in v0.4.0. setup.sh --register-project-dir writes a .curiosity/config.toml pointer at a non-code directory; scan.py then walks the pointer's configured [ingest] paths and ingests matching files into the workspace's vault. Threat shape: a malicious pointer file (committed to a project-dir, then synced to another user's machine) attempts to direct scan.py at filesystem paths outside the pointer's directory — /etc/passwd, ~/.ssh/, etc. — to exfiltrate sensitive content into the vault.
Mitigations:
code_repo.validate_pointer_paths()runs before any scan, rejecting paths with..segments, absolute paths, null bytes, or paths that resolve outside the pointer's directory (canonicalised viaPath.resolve()so symlinks can't paper over a..-free escape).scan.py oneandscan.py allrefuse to act on any pointer with unsafe paths and surface the violation. Invariant: a scan never reads bytes outside the pointer-dir's filesystem subtree.- Symlinks not followed by default.
[ingest] follow_symlinksdefaults tofalsein the pointer. Even when set totrue, scan.py rechecks every resolved path withrelative_to(pointer_dir)and skips anything that escaped — defence-in-depth against symlink-walking attacks. - Workspace itself cannot be a project-dir.
setup.sh --register-project-dirrefuses if the workspace path is at or under the proposed project-dir (would create an ingest loop scanning the vault's own outputs). - Extension whitelist enforced. The pointer's
[ingest] extensionsis a whitelist (default:.pdf, .md, .txt, .docx, .pptx, .csv, .xlsx, .html, .rst); nothing outside the list is read. No*glob escape..env,.ssh/, dotfiles, binaries, etc., are not on the default list. - Standard collateral always excluded.
.git/,.venv/,node_modules/,__pycache__/, etc., are added to the exclude list at scan time regardless of the pointer's user-supplied exclude entries. - Same untrusted-content envelope as today. Every extraction wraps body bytes in
<!-- BEGIN FETCHED CONTENT -->markers withuntrusted: truefrontmatter, runs throughscrub_check.py --mode vaultat write time, and quarantines on injection markers. No new write path; v0.4.0 just adds new INPUT paths to the existing pipeline.
The pointer file itself is committed (sharing routing across a team), so a malicious pointer in a shared project-dir would be visible in code review before it ran. Plus validate-paths is the failsafe regardless of code-review hygiene.
- Identifier resolution split.
scripts/identifier_cache.pyis cache-only: nourllibimport, no network calls. Workers append resolution requests to.curator/identifier-requests.jsonlinstead of triggering immediate lookups. scripts/identifier_resolve.pyis the only outbound-network script in the skill. It is:- Off by default (
identifier_resolution.enabled: falsein template config). - Bash-allowlist gated at the subcommand level. The allowlist permits only
identifier_resolve.py reviewandidentifier_resolve.py status(visibility, no network). Invokingidentifier_resolve.py run --yesrequires explicit user approval — there is no allowlist rule matching it. This is the load-bearing security gate; the config flag is convenience, not the boundary (see "Bash-allowlist as the boundary" below). - Endpoint-configurable. Defaults to PubChem/MyGene.info; enterprise users override
chemicals_endpoint/genes_endpointto point at internal mirrors. The resolver speaks PubChem-PUG-REST and MyGene.info shapes; mirrors must match those grammars.
- Off by default (
- The user sees exactly which names and endpoints are involved before approving any network call.
.curator/config.json is editable by orchestrator agents (the workspace's .claude/settings.json grants Edit(./.curator/**)). This is intentional — agents tune parallel_workers, switch model presets, adjust thresholds. But it means config flags cannot be the security boundary for sensitive operations: an agent that can write the flag can flip it.
The real security boundary is the bash allowlist itself — .claude/settings.json is not in the agent-writable scope (no Edit(./.claude/**) rule). Subcommand-level allowlist entries (e.g. Bash(uv run python3 .../identifier_resolve.py review:*) rather than the broader :*) let us permit visibility while requiring explicit user approval for state-changing operations.
When adding a sensitive operation:
- Don't gate it on a config flag alone.
- Pick subcommand-level allowlist entries that name the safe operations (
review,status, dry-run-style commands). - Leave the dangerous subcommand (the one that actually does the thing — egress, write, irreversible change) unallowed; the user is then prompted on every invocation.
- The config flag is fine as a default-off convenience switch on top, but it's belt; the allowlist is braces.
Currently in this scope:
identifier_resolve.py run— requires user approval (allowlist permits onlyreview/status).update.sh --yes— currently allowed under the broaderupdate.sh:*rule. Update is bounded by the hardcoded upstream slug (T6 mitigation), so the agent can't redirect; it can only trigger the upstream pull. Future hardening could narrow this further.
The skill makes outbound HTTP/HTTPS calls only at these well-defined sites. Auditors should grep for any urllib/requests/curl import or call outside this list:
| Site | Files | When | User control |
|---|---|---|---|
| GitHub (skill update) | scripts/update.sh |
update.sh --yes (explicit two-step) |
Hardcoded upstream slug; --source override validated against strict regex; non-default warned prominently. |
| PubChem PUG-REST / MyGene.info (or configured mirrors) | scripts/identifier_resolve.py |
identifier_resolve.py run --yes (explicit two-step) |
Off by default; enabled via config flag; endpoints configurable. |
| Astral uv installer | scripts/setup.sh |
None. Auto-install removed in v0.1.2. Setup refuses to install uv; prints instructions instead. | User runs the installer directly. |
| jsdelivr CDN (D3 / Fuse) | (formerly scripts/viewer.sh) |
None. Removed in v0.1.2. Vendored bundles ship in template/wiki-view/static/vendor/. |
Bundle bumps go through RELEASE_CHECKLIST.md review. |
Static analyzers will flag these patterns. Each call site is documented:
scripts/viewer_server.py:_rebuild—subprocess.run([list, of, hardcoded, args])invokingwiki_render.py. List-form argv, noshell=True, every argument hardcoded or derived fromPath(__file__).parent. No injection vector. Inline comment in the source.scripts/tables.py—re.compile(...)(regex compilation only). Earlier__import__("re").compile(...)form — same effect, replaced for static-analyzer clarity in v0.1.2.scripts/sweep.py— calls Python'sre.compile()extensively for pattern matching. Noeval, noexec, no dynamic code construction from external input.
The skill does not use eval, exec, compile() against externally-derived strings, or dynamic imports of externally-named modules anywhere in the trusted code path.
Found a security issue? Open an issue at https://github.com/benjsmith/curiosity-engine/issues, or for sensitive disclosures email the address listed in the GitHub profile of the maintainer (benjsmith). Skill is a one-person project; no formal CVE pipeline exists. Patch releases (vX.Y.Z) ship as fast as the fix can be tested and committed.