Turn research papers into visual, verifiable explanations.
PaperLens is a local-first research-paper ingestion and visualization platform in limited public beta. The application maintains strict separation between paper parsing, evidence provenance, semantic extraction, verification, and deterministic rendering stages.
backend/— FastAPI service and portable persistence boundaryfrontend/— Next.js reader shelldocs/— Architecture and incremental project state
Prerequisites: Python 3.11+, Node.js 20+, and npm. The backend uses SQLite by default for local development; PostgreSQL remains the production target behind the SQLAlchemy boundary.
python3 -m venv .venv
.venv/bin/pip install -r backend/requirements.txt
cp .env.example .env
.venv/bin/uvicorn app.main:app --app-dir backend --reloadHealth endpoints are available at http://localhost:8000/health/live and
http://localhost:8000/health/ready; /metrics exposes bounded local metrics.
Production configuration and migration guidance live in
docs/production.md, docs/operations.md,
and docs/security.md.
The API accepts a supported arXiv URL or identifier:
curl -X POST http://localhost:8000/api/papers/ingest \
-H 'Content-Type: application/json' \
-d '{"source":"https://arxiv.org/abs/1706.03762"}'After ingestion, the normalized source document and paragraph evidence are available through:
GET /api/papers/{paper_id}/document
GET /api/papers/{paper_id}/evidence/{evidence_id}
The document layer preserves source text, section/paragraph ordering, page numbers, and parser-provided coordinates. Phase 4 adds optional evidence-grounded research interpretation through the analysis endpoints below.
Phase 4 analysis endpoints are:
POST /api/papers/{paper_id}/extract
GET /api/papers/{paper_id}/analysis
Extraction is provider-neutral and evidence-grounded. Configure an OpenAI-compatible provider with the AI_* variables in .env.example. Without credentials, the API records safe FAILED/NO_EVIDENCE component states rather than fabricating analysis.
Phase 5 reader endpoints are:
GET /api/papers/{paper_id}/reader
GET /api/papers/{paper_id}/source
Open the visual reader at http://localhost:3000/papers/{paper_id}. It loads a reconstructed interactive paper when blocks exist (simplified explanation, inline diagrams, equations, figures, and tables in one scroll), fetches evidence passages only when requested, and opens the persisted PDF through the ownership-checked source endpoint. The earlier sectioned analysis remains under Source analysis.
Phase 6 verification endpoints are:
POST /api/papers/{paper_id}/verify
GET /api/papers/{paper_id}/verification
Verification is claim-level and persisted separately from PaperIR. It reports categorical support (SUPPORTED, PARTIALLY_SUPPORTED, UNSUPPORTED, CONTRADICTORY, UNVERIFIED) and never requires runtime AI to render the reader. Without verifier credentials, claims remain explicitly UNVERIFIED rather than receiving fabricated support.
Phase 7 Paper Chat endpoints are:
POST /api/papers/{paper_id}/chat/sessions
GET /api/papers/{paper_id}/chat/sessions/{session_id}
POST /api/papers/{paper_id}/chat/sessions/{session_id}/messages
Chat retrieves only bounded paragraph evidence from the current paper using deterministic BM25 ranking. Every assistant turn is persisted with document-bound citations; insufficient retrieval and missing AI credentials remain explicit safe states. CHAT_RETRIEVAL_TOP_K, CHAT_MAX_CONTEXT_CHARS, CHAT_MIN_RELEVANCE, and CHAT_MAX_QUESTION_CHARS tune the boundary.
Phase 8 research-intelligence endpoints are:
POST /api/workspaces
GET /api/workspaces
GET /api/workspaces/{workspace_id}
POST /api/workspaces/{workspace_id}/papers/{paper_id}
POST /api/workspaces/{workspace_id}/compare
GET /api/papers/{paper_id}/citation-graph
Phase 8 extracts source-first figure, table, equation, and reference artifacts into the Evidence Registry. RETRIEVAL_MODE supports LEXICAL, SEMANTIC, and HYBRID; provider-neutral semantic embeddings use a local SQLite cache and reciprocal-rank fusion when HYBRID_RETRIEVAL_ENABLED=true or hybrid mode is selected. For key-free local operation, set AI_PROVIDER=ollama with qwen3:4b and EMBEDDING_PROVIDER=ollama with nomic-embed-text; Ollama uses Metal when available and falls back to CPU. EMBEDDING_PROVIDER=hash remains available for deterministic tests only.
Relevant environment variables are documented in .env.example: database URL, arXiv timeout, local PDF storage path, PDF size limit, and frontend origin.
cd frontend
npm install
npm run devThe reader shell is available at http://localhost:3000.
The workspace shell is available at http://localhost:3000/workspaces.
Next.js is pinned to the validated 16.3.x line. The deterministic browser smoke suite uses Playwright with local API mocks:
npm run test:e2eThe research shell is available at http://localhost:3000/research. It creates a persistent, bounded run and reports observable planning, official arXiv discovery, candidate normalization/ranking, ingestion, extraction, synthesis, and verification events.
POST /api/research/runs
GET /api/research/runs/{run_id}
POST /api/research/runs/{run_id}/execute
POST /api/research/runs/{run_id}/cancel
GET /api/research/runs/{run_id}/report
Without credentials, planning and ranking remain deterministic, arXiv metadata is discovery-only, reports cite only persisted Evidence Registry tuples, and claims remain UNVERIFIED. Defaults are bounded by RESEARCH_MAX_SEARCH_QUERIES=6, RESEARCH_MAX_CANDIDATES=30, RESEARCH_MAX_INGESTED_PAPERS=8, RESEARCH_MAX_ITERATIONS=3, and ingestion concurrency 2.
Evaluation is offline-first and separate from production services. The versioned corpus manifest, annotation schemas, pure metrics, fixture runners, and report writer live under backend/evaluation/.
python -m backend.evaluation.run smoke --output-dir /tmp/paperlens-evaluation
python -m backend.evaluation.run all --output-dir /tmp/paperlens-evaluation
python -m backend.evaluation.run all --live --output-dir /tmp/paperlens-evaluationThe first two commands make no network or paid-provider calls. Reports record reproducibility metadata and remain PRELIMINARY until reviewed annotations and frozen production predictions are available. See docs/evaluation.md for corpus, metrics, failure taxonomy, and live/offline policy.
.venv/bin/python -m unittest discover -s backend/tests
.venv/bin/python -m compileall -q backend/app backend/testsWhen frontend dependencies are installed, also run:
cd frontend
npm run typecheck
npm run lint
npm run build
npm run test:e2eThe supported production-like architecture is a Next.js frontend behind an HTTPS edge, a FastAPI backend, PostgreSQL, and a durable paper-storage volume. There is no Redis, queue, or Kubernetes requirement at this scale. Application configuration stays environment-driven; provider-specific deployment files do not enter domain logic.
Build the local production-like stack with explicit PostgreSQL migrations:
docker compose up --build
python3 scripts/production-smoke.py
python3 scripts/load_test.py --base-url http://localhost:8000The image tags, migration revision, release version, and deployment timestamp should be recorded for every beta promotion. See docs/production.md, docs/release-checklist.md, and docs/beta-support.md for deployment, rollback, backup, retention, and support procedures.
Supported input is arXiv identifiers and URLs only. AI analysis, semantic
retrieval, Paper Chat, and the Research Agent are capability-gated; without
live credentials the reader, evidence, PDF, workspace, and BM25 paths remain
available while AI controls are disabled cleanly. Phase 10 evaluation results
are explicitly PRELIMINARY and are not claims of scientific accuracy.
Phase 12 — Deployment & Public Beta (v0.1.0-beta).
PaperLens also ships a local Tauri 2 macOS shell. It bundles the FastAPI
backend as a PyInstaller sidecar, the Next.js standalone server, and Node.js;
the app selects loopback ports at runtime and stores SQLite, PDFs, caches, and
logs under ~/Library/Application Support/com.paperlens.app/.
cd desktop && npm install
cd .. && ./script/build_and_run.sh --verifyThe release bundle is desktop/src-tauri/target/release/bundle/macos/PaperLens.app.
See docs/macos.md for lifecycle, token, data, signing, and
validation details.