Skip to content

About

Using Large Language Models to analyse neural circuit data, especially connectomics data

Resources

Stars

4 stars

Watchers

0 watching

Forks

Repository files navigation

Open in Colab

LLantia

LLM automated neural circuit inference and analysis

Most code here was generated by Claude under the instruction of the author(s).

LLM-powered pipeline for interpreting Drosophila neural circuits from connectome data. It first mines the literature for known cell-type functions, then uses that compendium — an example here — together with connectivity to infer pairwise pathway mechanisms, per-neuron integration logic, and a circuit-level synthesis.

Runs as SLURM array jobs on GPU cluster. Supports local inference (llama.cpp), OpenAI, and Anthropic.


Pipeline

paper_discovery/                  ← Step 0a: discover papers from seed DOIs/URLs
function_extraction_from_papers/  ← Step 0b: extract known cell-type functions from PDFs
neuron_interpretation/            ← Step 1: per-neuron functional hypotheses, grounded in Step 0's known functions
circuit_analysis/                 ← Step 2: 3-stage circuit analysis
embedding_analysis/               ← Optional: embed & explore outputs (UMAP, search)
llm_core/                         ← Shared LLM utilities

Setup

conda create -n llm python=3.12 && conda activate llm
pip install connectome_interpreter pandas scipy tqdm python-dotenv huggingface_hub \
            openai anthropic rapidfuzz PyMuPDF requests-cache
pip install -e .

API keys (skip for --provider llama):

echo "OPENAI_API_KEY=sk-..." >> .env
echo "ANTHROPIC_API_KEY=sk-ant-..." >> .env

For llama.cpp on HPC: docs/hpc_llama_setup.md.
--use-rag additionally requires chromadb tiktoken llama_index.


Usage

Before running: update paths at the top of each .sh file.
All scripts support --resume (skip completed entries) and --dump-prompts <dir> (write prompts to markdown for paste into Claude.ai / ChatGPT / Gemini).

Step 0a — Paper discovery (optional)

# Edit paper_discovery/seeds.txt (one DOI or URL per line)
sbatch paper_discovery/paper_discovery_hpc.sh
sbatch paper_discovery/match_and_download_pdfs.sh
# → paper_discovery/output/discovered_papers_with_local.csv

Step 0b — Function extraction (optional)

sbatch function_extraction_from_papers/function_extraction_from_paper.sh
sbatch function_extraction_from_papers/post_extraction.sh
# → function_extraction_from_papers/extraction_results/cell_function_summaries.csv

Step 1 — Neuron interpretation

sbatch neuron_interpretation/neuron_interpretation.sh
python neuron_interpretation/merge_chunks.py \
    --kind neuron --results-dir neuron_interpretation/results --output hypotheses.csv

Step 2 — Circuit analysis

python circuit_analysis/select_types.py        # → circuit_types.json
bash circuit_analysis/run_pipeline.sh          # submits steps 1→2→3 with SLURM deps
# or submit individually: run_step1.sh, run_step2.sh, run_step3.sh

Small circuits — one call instead of three stages:

sbatch circuit_analysis/run_small.sh           # → results_small/circuit_small.json

The 3-stage hierarchy exists because a raw all-to-all subgraph does not fit the context window of a large circuit. When it does fit, circuit_small.py asks step 3's question directly of the whole layered subgraph in a single call — same output schema, no decomposition. It estimates and prints the prompt size every run, and exits before loading the model if the prompt is over budget (--ctx-size minus --max-tokens, or --max-prompt-tokens); the prompt is saved either way, next to the result JSON.

Single-pair manual prompt: python circuit_analysis/dump_pair_prompt.py --source X --target Y --output prompt.md


LLM providers

Provider Default model Notes
llama Qwen3-35B-A3B GGUF Local inference; no rate limits; needs llama-server
openai gpt-4o Also gpt-5; Responses API with --reasoning-effort, --verbosity
anthropic claude-sonnet-4-6 Extended thinking via --thinking-budget

Override with --model <name>. Every script accepts --provider.


Key flags

Module Flag Default Description
circuit --side right Hemisphere (left / right)
circuit --n-steps 3 Max synaptic hops
circuit --threshold 0.01 Min connection weight
circuit --actor-critic off Critic-revision loop in step 1
circuit --shared-intermediates off Convergence analysis in step 2
neuron --n-steps 4 Connectivity hops
neuron --top-n 15 Top up/downstream types shown
neuron --no-add-function — Skip mapping cell types to known functions
extraction --mode experimental brief for short summaries
extraction --use-rag off ChromaDB RAG instead of full text

Parallelism

All long-running scripts accept --chunk-id / --n-chunks for SLURM arrays. Each chunk writes its own JSONL; merge with neuron_interpretation/merge_chunks.py (works for all steps).

About

Using Large Language Models to analyse neural circuit data, especially connectomics data

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages