A multi-agent question-answering system built on top of a Neo4j knowledge graph of NCU regulations.
This project extends a previous knowledge-graph QA pipeline by adding security validation, query planning, diagnosis, one-step repair, and explanation around the original KG retrieval flow. The goal is not only to answer a question, but also to detect unsafe requests and recover from failed or overly strict retrieval.
- 7-agent pipeline for regulation QA
- Read-only Neo4j / Cypher retrieval
- Security validation before KG access
- Diagnosis states:
SUCCESS,NO_DATA,QUERY_ERROR,SCHEMA_MISMATCH - One-round query repair with broader retrieval
- Hybrid answer generation:
- deterministic extraction for concise factual answers
- LLM fallback when direct extraction is not applicable
- Compatible with a fixed automated evaluation contract
| Metric | Result |
|---|---|
| Task Success Rate | 23.75 / 25 |
| Security & Validation | 15 / 15 |
| Error Detection Quality | 8 / 8 |
| Query Regeneration | 6 / 6 |
| Correct Resolution After Repair | 6 / 6 |
| System Performance | 58.75 / 60 |
Additional evaluation results:
- Unsafe request rejection: 10 / 10
- Failure-handling cases passed: 10 / 10
- Diagnosis label validity: 40 / 40
- Repair success when attempted: 8 / 8
- Normal QA accuracy improved from 15% to 95% after adding deterministic factual extraction before LLM fallback
flowchart TD
Q[User Question] --> NLU[1. NLU Agent<br/>Question → structured intent]
NLU --> SEC[2. Security Agent<br/>ALLOW / REJECT]
SEC -->|REJECT| RJ[Rejected Response<br/>No KG access]
SEC -->|ALLOW| PLAN[3. Planner Agent<br/>Build retrieval plan]
PLAN --> EXEC[4. Executor Agent<br/>A4 retrieval or read-only Cypher]
EXEC --> DIAG[5. Diagnosis Agent<br/>SUCCESS / NO_DATA / QUERY_ERROR / SCHEMA_MISMATCH]
DIAG -->|SUCCESS| ANS[Answer Generation]
DIAG -->|Non-SUCCESS| REP[6. Repair Agent<br/>Broaden / simplify query plan]
REP --> REXEC[Retry Executor<br/>One repair round only]
REXEC --> RDIAG[Re-diagnose Result]
RDIAG -->|SUCCESS| ANS
RDIAG -->|Still failed| FAIL[Failure / No evidence response]
ANS --> EXT{Deterministic extraction<br/>available?}
EXT -->|Yes| FACT[Concise factual answer]
EXT -->|No| LLM[LLM fallback]
FACT --> EXP[7. Explanation Agent]
LLM --> EXP
FAIL --> EXP
RJ --> EXP
EXP --> OUT[Final Output<br/>answer + safety_decision + diagnosis<br/>repair_attempted + repair_changed + explanation]
The pipeline uses a fixed front half — understand, validate, plan, execute, diagnose — and a conditional back half that triggers at most one repair attempt when retrieval fails.
| Article Nodes | Rule Nodes | Relationships |
|---|---|---|
![]() |
![]() |
![]() |
Transforms the raw question into structured intent.
It:
- extracts keyword variants
- classifies question type such as
penalty,time,yesno, orgeneral - detects the domain aspect such as
exam,admin, oracademic - marks underspecified questions as ambiguous
Runs before any KG access and rejects unsafe requests.
Checks include:
- dangerous database operations such as
DELETE,DROP, or write-oriented requests - possible Cypher injection patterns
- attempts to extract credentials, bulk records, or the entire KG
This keeps runtime QA read-only and prevents unsafe requests from reaching the executor.
Converts the structured intent into a retrieval plan and chooses between two first-pass behaviors.
- Clear questions use the proven A4 retrieval path (
get_relevant_articles) withuse_original=True. - Ambiguous / vague questions use direct Cypher with an intentionally unreachable
min_score=100, forcing aNO_DATAdiagnosis so the repair path can be exercised.
The planner also carries the detected aspect, question type, keywords, flags, and original question into downstream stages.
Supports two execution modes.
Mode 1 — A4 retrieval
For normal, clear questions, the executor reuses A4's get_relevant_articles(question) retrieval logic.
Mode 2 — direct read-only Cypher
For ambiguous first-pass queries and repaired plans, the executor runs a read-only Cypher query that:
- retrieves
Article→Rulecandidates - scores overlap across
Article.content,Rule.action,Rule.result, andRule.reg_name - applies a category bonus based on the detected aspect
- filters by
min_score - deduplicates the returned rules
Runtime QA remains read-only; the execution query does not perform KG writes.
Classifies execution into four states:
| State | Meaning |
|---|---|
SUCCESS |
Valid rows returned |
NO_DATA |
Query ran successfully but no useful rows matched |
QUERY_ERROR |
Runtime / connection failure |
SCHEMA_MISMATCH |
Error related to labels or properties |
The diagnosis result determines whether repair should run.
Runs only when the first attempt is not successful.
The repaired plan always switches to the direct-Cypher path (use_original=False) with a permissive threshold (min_score=1).
Repair behavior depends on the diagnosis:
SCHEMA_MISMATCH→ rebuild a small set of top-level match termsQUERY_ERROR→ simplify to the first few keywordsNO_DATA→ broaden keyword variants and rebuild match terms
Only one repair round is attempted per question.
Builds a short summary of what happened in the pipeline, including:
- question type / domain
- security decision
- diagnosis
- whether repair was attempted
- final answer preview
The A5 system reuses the KG built in the previous assignment without changing its schema.
graph LR
R[Regulation] -->|HAS_ARTICLE| A[Article]
A -->|CONTAINS_RULE| RULE[Rule]
Regulation
idnamecategory
Article
numbercontentreg_namecategory
Rule
rule_idtypeactionresultart_refreg_name
| Item | Count |
|---|---|
| Article nodes | 159 |
| Rule nodes | 199 |
CONTAINS_RULE relationships |
199 |
| Article coverage | 159 / 159 |
The main entry point is answer_question() in query_system_multiagent.py.
Simplified flow:
intent = nlu.run(question)
security = security_agent.run(question, intent)
if security["decision"] == "REJECT":
return rejected_result
plan = planner.run(intent)
execution = executor.run(plan)
diagnosis = diagnosis_agent.run(execution)
if diagnosis["label"] != "SUCCESS":
repaired_plan = repair_agent.run(diagnosis, plan, intent)
repaired_execution = executor.run(repaired_plan)
repaired_diagnosis = diagnosis_agent.run(repaired_execution)
if repaired_diagnosis["label"] == "SUCCESS":
execution = repaired_execution
diagnosis = repaired_diagnosis
if diagnosis["label"] == "SUCCESS":
answer = extract_fact_from_rules(...)
if answer is None:
answer = generate_answer(...)
return final_resultReturned object:
{
"answer": str,
"safety_decision": "ALLOW" | "REJECT",
"diagnosis": "SUCCESS" | "QUERY_ERROR" | "SCHEMA_MISMATCH" | "NO_DATA",
"repair_attempted": bool,
"repair_changed": bool,
"explanation": str,
}The implementation uses two first-pass paths rather than applying the same threshold to every question.
- Clear questions continue using the A4
get_relevant_articles()retriever. - Questions classified as vague or ambiguous are routed to direct Cypher with
min_score=100, intentionally producingNO_DATA. - Repair then switches to direct Cypher with
min_score=1and broader/simplified terms.
This preserves the working A4 behavior for normal QA while ensuring that the diagnosis → repair branch can actually be triggered and evaluated.
Security checks happen before Neo4j access.
This gives the pipeline a clear boundary:
unsafe request → REJECT → no KG query
safe request → ALLOW → continue pipeline
The original LLM-generated answers were often more verbose than the benchmark expected.
To improve factual QA, the system first tries deterministic extraction for short answers such as:
- scores
- fees
- durations
- yes/no outcomes
If no supported extraction rule applies, the system falls back to the LLM.
This improved normal-question accuracy from 15% to 95% in the provided evaluation.
The original A4 retrieval behavior was broad enough that almost every question returned some rows. As a result, diagnosis frequently reported SUCCESS, even for weak or intentionally vague test cases.
Fix: keep the A4 retriever for clear questions, but route ambiguous questions through a direct-Cypher path with an intentionally unreachable threshold (min_score=100). This guarantees a meaningful NO_DATA state, after which the repair agent retries with min_score=1.
Blocking only obvious database keywords was insufficient because unsafe intent can be written in normal language.
Fix: extended validation to cover write intent, bulk extraction, credential access, and injection-like keyword combinations.
The benchmark expects short factual outputs, while the LLM often generated unnecessary explanation.
Fix: added deterministic extraction before the LLM fallback.
- Python 3.11
- Neo4j
- Cypher
- Docker
- SQLite
- Hugging Face local LLM pipeline
- Regex / rule-based extraction
Assignment-5/
├── README.md
├── query_system_multiagent.py # main A5 pipeline entry point
├── agents/
│ └── a5_template.py # 7 agent implementations
├── auto_test_a5.py # automated evaluator
├── test_data_a5.json # benchmark cases
├── build_kg.py # KG builder inherited from A4
├── query_system.py # A4 retrieval / generation helpers
├── llm_loader.py # local model loader
├── ncu_regulations.db
├── requirements.txt
└── source/
docker start neo4jIf the container does not exist yet:
docker run -d --name neo4j \
-p 7474:7474 \
-p 7687:7687 \
-e NEO4J_AUTH=neo4j/password \
neo4j:latestpython -m venv venvWindows:
venv\Scripts\activatepip install -r requirements.txtpython build_kg.pypython auto_test_a5.pypython query_system_multiagent.pyThis project was developed against a fixed assignment benchmark. Some behaviors are intentionally benchmark-oriented rather than production-general:
- ambiguous questions are deliberately forced into
NO_DATAon the first pass to make the diagnosis / repair path observable - the concise factual-answer extractor contains question-specific rules and a few fixed fallbacks for known benchmark facts
For a production version, I would replace these benchmark-specific fallbacks with evidence-driven structured extraction, confidence checks, and more general semantic safety / repair logic.
This project helped me understand that a practical AI QA system needs more than a single LLM call.
The main lessons were:
- separate language understanding, retrieval, validation, and recovery responsibilities
- make failures explicit instead of hiding them behind weak retrieval results
- use deterministic logic when the expected output is structured and factual
- use LLM generation only where it adds value
- keep database access constrained and observable
The project also showed how an existing KG system can be extended without redesigning the underlying graph: most of the A5 work is in orchestration, validation, diagnosis, and repair around the A4 retrieval layer.







