Skip to content

Repository files navigation

atif-sql

CI security CodeQL Scorecard Python 3.13+ License Apache-2.0

ATIF-native analytics over agent trajectories.

Claude Code sessions (~/.claude/projects/**/*.jsonl) and Codex CLI rollouts (~/.codex/sessions/**/rollout-*.jsonl) are converted to ATIF (Harbor's Agent Trajectory Interchange Format, see the Harbor ATIF RFC: RFC-0001 in https://github.com/laude-institute/harbor), materialized as a corpus, and queried through DuckDB views. Converting once at the boundary — with an explicit, tested fidelity policy for what upstream drops — beats re-deriving trajectory semantics inside every SQL view. Both agents land in the same views, and sessions.agent says which one a row came from.

Install

One command, one package, every capability. Conversion, materialization, the DuckDB surface, the LLM analytics pipelines, and semantic search are all in the box — there are no extras to choose and nothing to install afterwards to make a command work.

uv tool install atif-sql     # the CLI on PATH
uvx atif-sql schema          # or run it without installing

Python 3.13 or newer. The install is substantial and deliberately so: 113 runtime dependencies, about 1.15 GiB on disk, because the analytics and vector paths carry polars, pyarrow, scipy, scikit-learn, umap-learn, hdbscan, lancedb, and duckdb. Prebuilt wheels cover CPython 3.13 on manylinux x86_64, macOS arm64, and Windows x86_64; Linux aarch64 compiles hdbscan from source, which needs a C toolchain. Alpine and other musl targets are not supported. RELEASING.md carries the measurements.

atif-sql analyze, atif-sql embed, and atif-sql search call Amazon Bedrock and cost money per invocation. Each one is dry-run by default and spends only when asked. Nothing else in the tool needs a credential.

How the source is organized

These seven directories under packages/ are internal structure, not seven installs. They exist so import-linter can enforce the layer and independence contracts at the source level; the only thing documented as installable is the atif-sql CLI above.

Directory What
atif-converter Claude Code and Codex CLI transcript → ATIF converters (ours, built on Harbor's public trajectory models) + per-agent fidelity policy (loss accounting per session)
atif-corpus Corpus materialization: discovery, watermarks, quiescence, atomic artifact writes
atif-duck DuckDB views + macros over the materialized corpus (core surface plus the v2 analytics surface)
atif-models Model alias registry + structured-output LLM client; no other package hardcodes a model id
atif-analytics eight v2 pipelines — five LLM (classify, trajectory, conflicts, friction, perceived) and three structural (cluster, terms, community)
atif-embed Cohere Embed v4 on Bedrock + LanceDB vector store + embedding backfill
atif-cli the composition root, and the source of the atif-sql command: convert, materialize, status, query, analyze, embed, search, examples, schema, cron

Quick start

Build the corpus from the local transcripts, then query it:

atif-sql materialize                   # discover sessions, convert, write the corpus
atif-sql status                        # corpus freshness, read-only
atif-sql query 'SELECT * FROM sessions LIMIT 5'

materialize also writes typed columnar artifacts (four parquet files per session) beside the JSON ones, so query parses no JSON for those sessions; --no-columnar skips them, older corpora keep working from trajectory.json, and status prints which path a corpus takes as query path.

For one session at a time, atif-sql convert <session.jsonl> converts and audits it in place.

materialize converts sessions across a process pool, min(8, cpu_count) workers by default. --workers N (or ATIF_SQL_MATERIALIZE_WORKERS) sets the size, and --workers 1 runs the single-process path. The pool doesn't change a byte of output: every worker writes the same artifacts through the same per-session staging directory and atomic rename.

Codex CLI transcripts

convert, materialize, and status all take --agent claude-code|codex, defaulting to claude-code. Pick codex and both roots move with it: the source becomes $CODEX_HOME (default ~/.codex) /sessions, and the corpus becomes ~/.atif-sql/corpus/codex. An explicit --source-root, --corpus-root, or the matching ATIF_SQL_* env var still wins.

atif-sql materialize --agent codex
atif-sql status --agent codex
atif-sql query "SELECT agent, count(*) FROM sessions GROUP BY 1"

One corpus holds one agent, so a Codex corpus and a Claude Code corpus stay separate directories, and one query reads one of them. Pass --corpus-root to pick which, or --agent codex to get the Codex default.

Working on atif-sql itself is a different setup — a clone, mise, and mise run check as the definition of done. CONTRIBUTING.md has it.

Agent workflow: schema → examples → query

atif-sql schema                    # every view + macro signature (<50 ms)
atif-sql examples                  # tested example queries, grouped core/analytics/vss
atif-sql query 'SELECT * FROM tool_rank(30) LIMIT 10'

atif-sql examples (or atif-sql query --examples) emits runnable queries derived from the catalog — not hardcoded strings — and every one is executed by the test suite against a fixture corpus, so the listing cannot rot. Piped output is JSON; filter with --requires core|analytics|vss and --category view|table-macro|scalar-macro.

Contributing

Setup, the five gates mise run check runs, the import-linter contracts you will trip, and the Conventional Commit rule the commit-msg hook enforces: CONTRIBUTING.md.

Releases are cut by commitizen and published to PyPI over OIDC Trusted Publishing — RELEASING.md covers the flow, the versioning model, and the install-weight measurements.

Security

Report a vulnerability privately through GitHub's advisory form, not in a public issue. Supported versions, the disclosure expectations, and what counts as a vulnerability in a tool that reads local transcripts are in SECURITY.md.

atif-sql query runs agent-composed SQL, so it runs it in a box: sized to the host before the corpus is registered (override with ATIF_SQL_QUERY_MEMORY_LIMIT and ATIF_SQL_QUERY_THREADS), a private spill directory outside the corpus that's removed on exit, no extension installs at query time (atif-sql embed --install-extension is where the lance extension comes from), file facing statements (COPY, EXPORT, ATTACH, INSTALL, LOAD, PREPARE, EXECUTE) refused before they run, and a refusal to run as root unless ATIF_SQL_ALLOW_ROOT=1 says so. The details and the one accepted disclosure (duckdb_settings() lists the granted paths) are in docs/reference/cli.md.

License

Apache License 2.0. Each of the seven module directories carries the same license file, so a published distribution ships it too.

About

ATIF-native analytics over Claude Code agent trajectories: convert sessions to ATIF, materialize a corpus, query it with DuckDB.

Topics

Resources

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages