Skip to content
View TarunSinghChauhan's full-sized avatar
🦇
🦇

Block or report TarunSinghChauhan

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
TarunSinghChauhan/README.md

TARUN SINGH CHAUHAN

AI Engineer — LLM Evaluation, RAG & Multi-Agent Systems



terminal

⚡ About

I build AI systems that go beyond API calls — evaluation frameworks, multi-agent pipelines, and RAG infrastructure that hold up under real testing, not just demos.

role:         AI Engineer
focus:        [LLM Evaluation, RAG Systems, Multi-Agent Orchestration]
education:    B.Tech, Artificial Intelligence & Data Science — Arya College of Engineering, Jaipur (2026)
status:       open_to_work — India / Remote / Global Sponsored

🛠️ Tech Stack



OpenAI Anthropic LangChain LangGraph Qdrant MLflow


🚀 Featured Projects

🔍 LLM Evaluation Harness & Red-Teaming

Benchmarks GPT-4o & Claude Sonnet across 50 MMLU prompts (ROUGE-L, BERTScore, bootstrapped 95% CI). LLM-as-judge ensemble — Cohen's κ = 0.81. 20+ adversarial red-team patterns. Regression detection: days → under 4 minutes.

Python LangGraph LangSmith MLflow FastAPI

→ View Repo

🤖 Multi-Agent Code Review System

3-agent LangGraph pipeline — static analysis → OWASP scanner → LLM fix proposal. Caught 4 critical vulnerabilities in benchmark testing. $0.50 per-review cost budget enforced.

Python LangGraph OpenRouter PostgreSQL

→ View Repo

📈 Autonomous Financial Research Agent

4-step reasoning chain (Market Context → Analysis → Risk → Thesis) via Yahoo Finance + DuckDuckGo APIs. SHA-256 reproducibility hash per report. ~60% cost reduction via caching. Reports in under 60 seconds.

Python FastAPI OpenRouter Redis

→ View Repo

🎬 CodePulse — Cinematic Code Walkthrough

Sends code to LLaMA 4 via Groq, returns a structured ExecutionScript JSON driving a live animated walkthrough — variable tracking, call stack, voice narration, "Break Mode" bug injection.

Next.js 14 TypeScript Groq API

→ View Repo

🧩 Multimodal Product Intelligence Platform

End-to-end multimodal pipeline — embeddings, vector search, and LLM inference served through a Streamlit interface, fully containerized.

Python Vector Search Streamlit Docker

→ View Repo


💼 Experience

Data Science and AI Intern · ZIDIO Development — Remote · Jun–Aug 2025

Built AI-powered Python automation pipelines integrating LLM APIs across 5+ sources, reducing downstream data errors by ~35–40%. Shipped prompt engineering workflows and Power BI dashboards saving ~4–5 hrs/week. Wrote pytest coverage and structured logging across production pipelines.

Generative AI & ML Intern · Linux World Pvt. Ltd. — Jaipur · Jul–Aug 2024

Engineered ML data pipelines and applied GenAI techniques across two projects — an NLP chatbot on LLM APIs and a regression-based price predictor. Preprocessing and feature engineering in Pandas/NumPy.


🦇 Contribution Signal (real activity, auto-updates daily)

bat contribution matrix

🎯 Highlights

Eval Agents RAG Shipped


📫 tarunsinghchauhan088@gmail.com  |  💼 Open to AI Engineer / GenAI Engineer / LLM Engineer roles — India, Remote & Global Sponsored

Pinned Loading

  1. multi-agent-code-review multi-agent-code-review Public

    3-agent LangGraph pipeline for automated code review — static analysis, OWASP security scanning, and LLM-generated fix proposals with enforced cost budgets.

    Python 2 1

  2. CodePulse CodePulse Public

    Full-stack AI app that turns code into a cinematic, narrated walkthrough. Sends code to LLaMA 4 via Groq, returns a structured ExecutionScript JSON driving live variable tracking, call-stack visual…

    TypeScript 1

  3. finance-research-agent finance-research-agent Public

    Tool-augmented autonomous agent running a 4-step reasoning chain (Market Context → Financial Analysis → Risk Assessment → Investment Thesis) via Yahoo Finance and DuckDuckGo APIs. SHA-256 reproduci…

    Python 1

  4. Multimodal-Product-Intelligence-Platform Multimodal-Product-Intelligence-Platform Public

    End-to-end multimodal product intelligence pipeline — embeddings, vector search, and LLM inference served through a Streamlit interface, fully containerized.

    Python 1

  5. llm-eval-harness llm-eval-harness Public

    Production-grade LLM evaluation framework benchmarking GPT-4o and Claude Sonnet with ROUGE-L, BERTScore, and LLM-as-judge red-teaming for regression and safety detection.

    Python 1

  6. ai-engineering-projects ai-engineering-projects Public

    "Curated index of production-grade AI/LLM engineering projects — evaluation, agents, RAG, and cost-bounded systems." Also add topics/tags: llm, langchain, langgraph, rag, ai-agents, fastapi — this …

    1