I build practical AI systems that combine LLMs, deterministic business rules, retrieval, tool use, human review, evals, and production-style engineering to solve real-world business and customer problems.
Python AWS Bedrock LangGraph FastAPI Gradio PostgreSQL pgvector ChromaDB Neo4j MCP Docker Pydantic
Grounded appliance repair assistant — cited ask, multi-turn diagnose, safety gates, and opt-in LLM semantic curation over manufacturer service docs
Appliance repair needs authoritative manufacturer knowledge, not generic LLM advice. This framework turns service manuals and tech sheets into a grounded ask + multi-turn diagnose product: numbered citations, model/serial applicability filters, and deterministic safety policy — proven end-to-end on a Whirlpool WFW5620H reference corpus.
- Experiment-driven pipeline: parser bake-offs → hybrid layout routing → structured chunking → hybrid retrieval → grounded Q&A and LangGraph diagnose, each layer gated by manual eval benches and ADRs
- Corpus curator board (
/ui/corpus): propose → finalize → review, with assist that only suggests — saving and activating a corpus version stays human-owned - Applicability and document precedence as structured manifest data — wrong-model docs are filtered before generation, not left to embedding similarity alone
- Deterministic safety and grounding: allow / warn / escalate / block for owner vs technician; pre-/post-LLM gates; optional weak-evidence abstain so off-topic packs never reach the answer LLM
- Product surface: FastAPI + streaming web UI (ask, diagnose, search, session path tally); optional Langfuse traces and
mine-tracesto draft evals from live failures - 662 automated tests + layer benches (parsing, retrieval, QA, safety, chain smoke)
Why native PDF chunks for high-risk docs: spend a capable LLM once at corpus build so procedures keep warnings with their steps, index compact representations for retrieval, then attach native PDF page-ranges (or rasters) at generate time so the model sees structure, layout, and spatial cues that text-only chunks lose. A human curator approves the markers on the real PDF before cutover — a one-time cost per doc version that's easy to justify on consequential material like repair procedures, where a wrong step is a safety or warranty event, not just a bad answer.
Python · LangGraph · FastAPI · PostgreSQL · pgvector · BGE · Langfuse · OpenAI · RAG
AI customer support agent for a DTC electronics retailer, built on AWS Bedrock
Classifies support requests, routes each task type to the lowest-viable Bedrock model, answers via RAG + Lambda tools over order/return/refund data, and escalates to a human for refunds over $100 or low-confidence cases — with quality and cost claims backed by versioned eval scorecards, not assumptions.
- API Gateway → LangGraph on Lambda; Bedrock Converse, Guardrails, and Knowledge Bases over S3 Vectors
- DynamoDB routing table driven by offline scorecards → policy generate/adopt → publish (lowest-viable model per task type)
- Isolated tool Lambdas (order / return / refund) with least-privilege IAM; refunds over $100 require human approval
- Separate eval plane: golden scenarios, live Bedrock candidates, LLM-as-judge, prompt-cache measurement, GitHub-first ADRs/releases
- Delivered via a GitHub-first SDLC with build-time roles (architect, implementation planner, code reviewer) served over MCP by enterprise-sdlc-mcp — kept separate from the runtime agent
Python · AWS · Bedrock · LangGraph · CDK · DynamoDB · MCP
End-to-end college admissions planner with hybrid scoring and grounded AI
Helps a family organize a high-stress planning process — school research, shortlist, fit/cost cues, applications, and a final decision — in one full-stack workspace instead of spreadsheets and one-off AI chats. AI is a core capability (research, compare, coaching), but admit odds stay on a deterministic scoring floor with labeled estimates, provenance, and confidence rather than vibes-based answers.
- Shared College Knowledge Base (Scorecard + Common Data Set + LLM research) with field provenance; separate student workspaces for shortlists, assessments, and trackers
- Hybrid admit %: rule-table primary estimate + factor breakdown; LLM explains and may offer a secondary opinion without overriding the planning %
- Guided journey from profile → discover → assess/compare → apply/track → decide, ending in a decision table backed by a fact matrix and an AI narrative grounded in the research package
- GitHub-first delivery with durable docs/ADRs, CI, and Playwright smoke tests
Python · FastAPI · Next.js · PostgreSQL · pgvector · OpenAI · MCP
Reusable build-time SDLC agent catalog, served over MCP
A shared, versioned definition of how AI coding agents deliver software — one catalog of agent roles (Product Analyst, Solution Architect, Code Reviewer, etc.) and review checklists, installed per-project and served over MCP, so every repo doesn't reinvent its own review bar or quietly let it drift.
- 9 agent roles + 31 skills spanning generic SDLC checklists (PR review, architecture, application security, supply chain, CI/CD, postmortems) and stack-specific ones (IAM least privilege, Bedrock Guardrails, knowledge-graph modeling, Graph RAG retrieval)
- Agents carry machine-readable code-modify permissions and path allowlists, and skills declare when they apply — guardrails the tooling enforces, not prose a human has to trust
- Copy-in project scaffold (manifest, docs skeleton, AGENTS.md, CI) so a new repo meets the catalog's contract on day one instead of reverse-engineering it from an existing consumer
- PR review enforces Blocker/Major/Minor severity with cited file+line evidence, gated by a Cursor hook so the review actually happens before merge
Python · MCP · SDLC Tooling
Agentic assistant for customer service order investigation and resolution
Investigates delayed, backordered, and exception-based orders by combining deterministic evidence gathering, policy-grounded RAG, and bounded agentic reasoning — with human approval required before any customer-facing recommendation.
- LangGraph investigation workflow over synthetic ERP data (Postgres)
- pgvector retrieval for policies and similar past cases
- Deterministic feasibility engine + human-in-the-loop review gate
- MCP tools for order/customer/inventory/policy lookup, shared across FastAPI + Gradio
- Golden eval cases for root cause, citation validity, and action acceptability
Python · LangGraph · pgvector · MCP · FastAPI
Multi-agent clinical workflow demonstrator (synthetic data only)
Specialist agents debate a diagnosis under LangGraph orchestration; deterministic risk gates flag high-stakes cases and route them to a human reviewer before any conclusion stands.
- Multi-agent routing with bounded debate cycles
- Citation-backed RAG over guideline-style documents
- Deterministic safety/risk checks independent of the LLM
Python · LangGraph · Postgres · Chroma · MCP
Enterprise RAG-based data catalog and lineage assistant
Semantic table discovery, lineage exploration, change-impact analysis, and validated natural-language-to-SQL — served from one core service through three interfaces.
- Dual-store architecture: Postgres (structured) + ChromaDB (semantic)
- RAG-scoped SQL generation with validation before execution
- Exposed via Gradio UI, FastAPI REST API, and a protocol-compliant MCP server
Python · RAG · MCP · FastAPI · Gradio
Knowledge-graph dependency intelligence — evidence-backed impact, ownership, and change risk over a synthetic enterprise
Large enterprises scatter dependency knowledge across CMDBs, API catalogs, and tribal memory. This app unifies synthetic Meridian Retail Group catalogs into a Neo4j knowledge graph so architects and engineers can search systems, traverse multi-hop dependencies, ask grounded NL questions, run bounded investigations, and score change risk — with every claim tied to provenance, not vibes.
- Deterministic explorer (search, paths, ownership/capability rollups, interactive Cytoscape graph) plus closed NL templates with zero LLM on the core path
- Optional Graph RAG / Hybrid Doc RAG Ask and bounded Agentic Investigate behind feature flags — answers cite retrieved graph/doc evidence or refuse
- Advanced Intelligence: change-risk scoring, read-only retirement what-if, drift signals, and technology rationalization
- Idempotent ingestion with entity resolution and dry-run source–graph reconciliation; golden evals + GitHub Actions CI; ADRs and architecture diagrams in-repo
Python · Neo4j · FastAPI · HTMX · Graph RAG · RAG
Healthcare EDI validation and synthetic encounter data toolkit
A Python toolkit for validating healthcare encounter files and generating realistic synthetic claims data inspired by payer/provider workflows.
- Synthetic claims, provider, and member data generation
- 837 encounter validation, 999/835-style response concepts
- Rule-based validation layers with traceable output
Python · EDI · Healthcare Data
- Deterministic floor, agentic ceiling — business-critical facts and constraints come from deterministic systems; agents reason within those boundaries, not around them.
- Evidence before recommendation — every AI output is grounded in retrieved policies, structured records, or traceable source documents, with provenance recorded down to the source system or document version, not just a citation label.
- Traceable decisions, not just traceable facts — every run records the conditions behind its action: which tool ran with which arguments, which gate or threshold fired, which model was routed, and why it scored, escalated, or abstained — as a bounded step trace, not a free-form agent diary.
- Refusal is a feature — when retrieved evidence is weak or missing, the system abstains or says what it can't support instead of generating a confident-sounding guess.
- Human review for consequential actions — for customer, clinical, or financial decisions, AI assists; it doesn't silently execute.
- Human effort spent where it compounds — one-time curation at build time (approving document structure, routing policy, graph mappings) beats re-reviewing every answer at runtime, and pays for itself on consequential material.
- Evals are part of the product — golden cases, citation checks, and action-validity checks ship alongside the feature, not after it.
- Tools reusable across interfaces — the same capability is callable from UI, API, and MCP clients, not locked inside one demo screen.
- Local-first demos, production-patterned architecture — even synthetic-data projects get real tests, CI, and documented operating assumptions.
- The delivery process is engineered too — agent roles, review checklists, and ADRs are versioned artifacts with machine-readable permissions, so AI-assisted work is auditable the same way the code is.
- Cost is a design constraint — route each task to the lowest-viable model, not the biggest one, and default cloud infrastructure to a near-zero-cost idle state when not in active use.
- GitHub: raghuram-chittibomma
- LinkedIn: https://www.linkedin.com/in/raghuram-chittibomma-3312252/
- Email: raghuram.chittibomma@gmail.com



