A research project that compares two approaches to AI question answering: Retrieval-Augmented Generation (RAG) and Knowledge-Augmented Generation (KAG).
The project builds both pipelines from scratch, runs them against the same questions under identical conditions, and evaluates the results using an automated LLM-as-Judge scoring system. It ships with a web dashboard that visualises benchmark results and lets you run live comparisons interactively.
RAG has become the standard approach for grounding LLM responses in external knowledge. You embed your documents, retrieve the most relevant chunks for a given question, and feed them to the model as context. This works well for straightforward factual lookups where the answer sits inside a single text chunk.
But many real questions require connecting facts that live in different documents. For example: "The AI that beat the world Go champion was created by which lab, and what large language model did that lab later build?" Answering this requires chaining AlphaGo to DeepMind to Gemini, a path that crosses document boundaries. RAG retrieves text chunks independently and has no mechanism to follow these relationship chains.
KAG addresses this by constructing a knowledge graph during ingestion. Named entities are extracted from each document, relationships between them are identified, and everything is stored as a graph in Neo4j. At query time, KAG first links entities in the question to graph nodes, traverses their relationships to gather structured context, and merges this with the standard vector-retrieved chunks before prompting the LLM.
This project provides the infrastructure to build both pipelines, measure where each one excels, and demonstrate the differences with real queries and quantitative metrics.
After running the setup and evaluation, you will have:
- A Neo4j knowledge graph built from the sample documents, with entities and relationships you can explore in the Neo4j Browser.
- An evaluation report comparing RAG and KAG across six multi-hop test queries, scored on factual accuracy, hallucination rate, context relevance, and latency.
- A web dashboard at
http://localhost:8000with four sections:- Project overview and architecture diagrams
- Evaluation methodology explanation
- Benchmark results with per-query breakdowns
- An interactive playground where you can ask your own questions and see both pipelines respond in real time, with on-the-fly LLM-as-Judge scoring
The current benchmark results (on the included sample documents) show:
| Metric | RAG | KAG | Winner |
|---|---|---|---|
| Factual Accuracy | 67.4% | 84.2% | KAG |
| Hallucination Rate | 34.7% | 15.8% | KAG |
| Context Relevance | 51.2% | 51.2% | Tied |
| Avg Latency | 2993ms | 1691ms | KAG |
KAG shows the largest advantage on queries that require reasoning across multiple documents, which is the core hypothesis this project tests.
Before running anything, make sure the following are installed and accessible:
- Python 3.10 or higher
- Neo4j Desktop (or Neo4j Community Server)
- Download from neo4j.com/download
- Create a local database and set a password
- Start the database so it listens on
bolt://localhost:7687
- An NVIDIA NIM API key
- Sign up at build.nvidia.com
- Generate an API key from the dashboard
- SpaCy English model (installed automatically during setup, or manually with
python -m spacy download en_core_web_sm)
Follow these steps in order. Each step depends on the previous one completing successfully.
git clone https://github.com/your-username/rag_kag.git
cd rag_kag
python -m venv .venv
# On Windows (PowerShell)
.\.venv\Scripts\Activate.ps1
# On macOS/Linux
source .venv/bin/activate
pip install -r backend/requirements.txt
python -m spacy download en_core_web_smCopy the example file and fill in your credentials:
cp .env.example .envOpen .env and set:
NVIDIA_API_KEY=nvapi-your-key-here
NEO4J_URI=bolt://localhost:7687
NEO4J_USER=neo4j
NEO4J_PASSWORD=your-neo4j-password
The remaining defaults (CHROMA_PERSIST_DIR, LLM_MODEL, EMBEDDING_MODEL,
CHUNK_SIZE, CHUNK_OVERLAP) work out of the box.
Open Neo4j Desktop, start your local database, and verify it is running. You should
be able to access the Neo4j Browser at http://localhost:7474.
This single command will:
- Verify API connectivity (NVIDIA NIM and Neo4j)
- Load and chunk the sample documents
- Build embeddings and store them in ChromaDB
- Extract entities and relationships via LLM and build the Neo4j knowledge graph
- Run all six test queries through both RAG and KAG pipelines
- Score every response with the LLM-as-Judge evaluator
- Save results to
data/evaluation_results.json
python test_e2e.pyThe first run takes approximately 3 to 5 minutes because it builds the knowledge graph from scratch. If you need to re-run just the queries and evaluation (skipping ingestion), use:
python test_e2e.py --skip-ingestuvicorn backend.main:app --host 0.0.0.0 --port 8000 --reloadThen open your browser to http://localhost:8000.
- Sections 1-3 of the dashboard display the project explanation, methodology, and benchmark results loaded from the evaluation file.
- Section 4 lets you type a question, run both pipelines live, and see real-time typewriter-style answers with evaluation scores.
- Open the Neo4j Browser at
http://localhost:7474and runMATCH (n:Entity)-[r:RELATION]->(m:Entity) RETURN n,r,m LIMIT 150to visualise the knowledge graph.
rag_kag/
backend/
config.py # Centralised settings from .env
main.py # FastAPI app, lifespan, static file serving
ingestion/
loader.py # Document loading (txt, pdf)
chunker.py # Text splitting with overlap
embedder.py # Sentence-transformer embeddings
vector_store.py # ChromaDB persistence
graph_builder.py # LLM-based triple extraction, Neo4j writer
rag/
retriever.py # ChromaDB vector search
prompt_builder.py # RAG prompt construction
pipeline.py # Full RAG pipeline orchestration
kag/
entity_linker.py # SpaCy NER + Neo4j fuzzy matching
graph_traversal.py # Multi-hop graph path retrieval
context_merger.py # Combines graph triples with vector chunks
pipeline.py # Full KAG pipeline orchestration
evaluation/
__init__.py # Evaluation module (LLM-as-Judge used in routes)
routes/
ingest.py # POST /api/ingest
query.py # POST /api/query
graph.py # GET /api/graph/*
metrics.py # GET /api/metrics, POST /api/evaluate, POST /api/judge
frontend/
index.html # Single-page dashboard
style.css # Design system and styles
app.js # Client-side application logic
data/
sample_docs/ # Source documents (5 AI-related texts)
evaluation_results.json # Generated benchmark results
test_e2e.py # End-to-end test and evaluation runner
.env.example # Template for environment configuration
.gitignore
The following companion documents provide additional detail on specific topics:
| Document | Contents |
|---|---|
| ARCHITECTURE.md | Technical deep dive into the RAG and KAG pipelines, system design, and technology stack |
| BENCHMARKS.md | Detailed evaluation results, per-query analysis, and guidance on interpreting the metrics |
| TESTING.md | How to run automated tests, manual verification steps, and API smoke tests |
| SECURITY.md | Credential handling, API key management, and deployment considerations |
This project is provided for educational and research purposes.