LLM automated neural circuit inference and analysis
Most code here was generated by Claude under the instruction of the author(s).
LLM-powered pipeline for interpreting Drosophila neural circuits from connectome data. It first mines the literature for known cell-type functions, then uses that compendium — an example here — together with connectivity to infer pairwise pathway mechanisms, per-neuron integration logic, and a circuit-level synthesis.
Runs as SLURM array jobs on GPU cluster. Supports local inference (llama.cpp), OpenAI, and Anthropic.
paper_discovery/ ← Step 0a: discover papers from seed DOIs/URLs
function_extraction_from_papers/ ← Step 0b: extract known cell-type functions from PDFs
neuron_interpretation/ ← Step 1: per-neuron functional hypotheses, grounded in Step 0's known functions
circuit_analysis/ ← Step 2: 3-stage circuit analysis
embedding_analysis/ ← Optional: embed & explore outputs (UMAP, search)
llm_core/ ← Shared LLM utilities
conda create -n llm python=3.12 && conda activate llm
pip install connectome_interpreter pandas scipy tqdm python-dotenv huggingface_hub \
openai anthropic rapidfuzz PyMuPDF requests-cache
pip install -e .API keys (skip for --provider llama):
echo "OPENAI_API_KEY=sk-..." >> .env
echo "ANTHROPIC_API_KEY=sk-ant-..." >> .envFor llama.cpp on HPC: docs/hpc_llama_setup.md.
--use-rag additionally requires chromadb tiktoken llama_index.
Before running: update paths at the top of each .sh file.
All scripts support --resume (skip completed entries) and --dump-prompts <dir> (write prompts to markdown for paste into Claude.ai / ChatGPT / Gemini).
# Edit paper_discovery/seeds.txt (one DOI or URL per line)
sbatch paper_discovery/paper_discovery_hpc.sh
sbatch paper_discovery/match_and_download_pdfs.sh
# → paper_discovery/output/discovered_papers_with_local.csvsbatch function_extraction_from_papers/function_extraction_from_paper.sh
sbatch function_extraction_from_papers/post_extraction.sh
# → function_extraction_from_papers/extraction_results/cell_function_summaries.csvsbatch neuron_interpretation/neuron_interpretation.sh
python neuron_interpretation/merge_chunks.py \
--kind neuron --results-dir neuron_interpretation/results --output hypotheses.csvpython circuit_analysis/select_types.py # → circuit_types.json
bash circuit_analysis/run_pipeline.sh # submits steps 1→2→3 with SLURM deps
# or submit individually: run_step1.sh, run_step2.sh, run_step3.shSmall circuits — one call instead of three stages:
sbatch circuit_analysis/run_small.sh # → results_small/circuit_small.jsonThe 3-stage hierarchy exists because a raw all-to-all subgraph does not fit the context
window of a large circuit. When it does fit, circuit_small.py asks step 3's question
directly of the whole layered subgraph in a single call — same output schema, no
decomposition. It estimates and prints the prompt size every run, and exits before
loading the model if the prompt is over budget (--ctx-size minus --max-tokens, or
--max-prompt-tokens); the prompt is saved either way, next to the result JSON.
Single-pair manual prompt: python circuit_analysis/dump_pair_prompt.py --source X --target Y --output prompt.md
| Provider | Default model | Notes |
|---|---|---|
llama |
Qwen3-35B-A3B GGUF | Local inference; no rate limits; needs llama-server |
openai |
gpt-4o |
Also gpt-5; Responses API with --reasoning-effort, --verbosity |
anthropic |
claude-sonnet-4-6 |
Extended thinking via --thinking-budget |
Override with --model <name>. Every script accepts --provider.
| Module | Flag | Default | Description |
|---|---|---|---|
| circuit | --side |
right |
Hemisphere (left / right) |
| circuit | --n-steps |
3 |
Max synaptic hops |
| circuit | --threshold |
0.01 |
Min connection weight |
| circuit | --actor-critic |
off | Critic-revision loop in step 1 |
| circuit | --shared-intermediates |
off | Convergence analysis in step 2 |
| neuron | --n-steps |
4 |
Connectivity hops |
| neuron | --top-n |
15 |
Top up/downstream types shown |
| neuron | --no-add-function |
— | Skip mapping cell types to known functions |
| extraction | --mode |
experimental |
brief for short summaries |
| extraction | --use-rag |
off | ChromaDB RAG instead of full text |
All long-running scripts accept --chunk-id / --n-chunks for SLURM arrays. Each chunk writes its own JSONL; merge with neuron_interpretation/merge_chunks.py (works for all steps).