Refract is a session-centered research workspace for literature review and statistical analysis. It keeps papers, saved evidence, datasets, and analysis runs inside one continuous environment so qualitative reading and quantitative modeling do not have to be reconstructed across separate tools.
The current implementation is built around an Industrial and Systems Engineering workflow: upload papers into a research session, extract and organize evidence, compare sources against a live research goal, attach a dataset to the same session, and run a staged analysis flow that surfaces profile, audit, model, and interpretation outputs with explicit provenance.
Applied research work is usually split across PDF readers, note-taking tools, spreadsheets, notebooks, and writing documents. That separation creates friction:
- reading notes drift away from the passages that produced them
- cross-paper synthesis gets rebuilt from scattered fragments
- statistical interpretation is often written after the fact, without the literature context that should shape it
Refract addresses that by making the research session the continuity layer. The same session holds the papers being read, the evidence saved from those papers, the dataset being analyzed, and the run history used to interpret the results.
Reader: upload PDFs, read them in-app, select passages, and generate passage-grounded responses from the active document.Index: turn saved questions, notes, concept tags, and source-linked responses into persistent research memory.Compare: build structured cross-paper synthesis through matrices, topic coverage, and gap analysis tied to the session goal.Review: surface study queues and review artifacts built from saved evidence.Stats: attach a session-scoped dataset and run a staged quantitative workflow for profiling, audit, modeling, and interpretation.Research session backbone: keep the qualitative and quantitative tracks connected through one shared session object rather than separate project fragments.
The live implementation is organized around a few concrete seams:
- frontend/src/App.jsx: workspace shell that exposes
Reader,Index,Compare,Review, andStats. - backend/app/main.py: FastAPI entrypoint that wires PDF, chat, highlight, review, research-session, data-file, and analysis routes.
- backend/app/models/research_session.py: session model that coordinates papers, datasets, and analysis runs.
- backend/app/services/comparative_analysis_service.py: cross-paper synthesis layer for session-aware compare output.
- backend/app/services/analysis_execution_service.py: staged analysis execution, including profiling and audit flow.
- backend/app/services/analysis_models_service.py: regression-model benchmarking and model output generation.
- backend/app/services/analysis_evidence_service.py: evidence provenance labels for statistical interpretation.
At runtime, the app uses React/Vite in the frontend, FastAPI in the backend, PostgreSQL for relational state, ChromaDB for document-vector storage, and object storage for uploaded files and generated artifacts.
The primary end-to-end quantitative demonstration in the current project is Battery Remaining Useful Life prediction. The Stats workspace is designed to keep that analysis inside the same session as the supporting literature, so the interpretation layer can reconnect modeling decisions to the papers already in scope.
This repository is best understood as an active research prototype rather than a polished general-purpose analytics platform. The strongest implemented thread is the integrated workflow itself: literature-grounded reading, persistent indexing, structured comparative synthesis, and session-scoped quantitative analysis.
- Node.js for the frontend dev server
- A Python environment that satisfies backend/requirements.txt
- PostgreSQL running on
localhost:5432 - ChromaDB available on
localhost:8001 - A populated
backend/.envfile
- Copy the example environment file and fill in the required values:
cp backend/.env.example backend/.env- Start the local stack from the repo root:
./start.sh- Open the app at http://localhost:5173.
Useful helper scripts:
./status.shchecks PostgreSQL, ChromaDB, backend, and frontend status../stop.shstops the local services started bystart.sh.
If you prefer containers, the repository also includes docker-compose.yml:
cp backend/.env.example backend/.env
docker-compose up --buildThis brings up:
- frontend on
http://localhost - backend on
http://localhost:8000 - PostgreSQL on
:5432 - ChromaDB on
:8001
The backend environment file currently expects values for:
ANTHROPIC_API_KEYOPENAI_API_KEYTAVILY_API_KEYAWS_ACCESS_KEY_IDAWS_SECRET_ACCESS_KEYS3_BUCKET_NAMEDATABASE_URL
Full functionality depends on those services being configured, especially for file storage and AI-assisted response paths.
frontend/ React + Vite workspace shell and UI components
backend/ FastAPI app, models, routers, and research services
docs/ planning notes, implementation notes, and README assets
playwright-tests/ browser probes and workflow verification scripts
research/ project notes, literature framing, and decision logs
scripts/ small project-specific utilities
- Some internal file names and scripts still use the earlier working label
pdf-workspace; the product name isRefract. - The screenshots and diagrams in this README reflect the current research workflow and Battery RUL demonstration used in the project deep-dive material.





