I build AI systems that go beyond API calls — evaluation frameworks, multi-agent pipelines, and RAG infrastructure that hold up under real testing, not just demos.
role: AI Engineer
focus: [LLM Evaluation, RAG Systems, Multi-Agent Orchestration]
education: B.Tech, Artificial Intelligence & Data Science — Arya College of Engineering, Jaipur (2026)
status: open_to_work — India / Remote / Global Sponsored|
Benchmarks GPT-4o & Claude Sonnet across 50 MMLU prompts (ROUGE-L, BERTScore, bootstrapped 95% CI). LLM-as-judge ensemble — Cohen's κ = 0.81. 20+ adversarial red-team patterns. Regression detection: days → under 4 minutes.
|
3-agent LangGraph pipeline — static analysis → OWASP scanner → LLM fix proposal. Caught 4 critical vulnerabilities in benchmark testing. $0.50 per-review cost budget enforced.
|
|
4-step reasoning chain (Market Context → Analysis → Risk → Thesis) via Yahoo Finance + DuckDuckGo APIs. SHA-256 reproducibility hash per report. ~60% cost reduction via caching. Reports in under 60 seconds.
|
Sends code to LLaMA 4 via Groq, returns a structured
|
|
End-to-end multimodal pipeline — embeddings, vector search, and LLM inference served through a Streamlit interface, fully containerized.
|
|
Data Science and AI Intern · ZIDIO Development — Remote · Jun–Aug 2025
Built AI-powered Python automation pipelines integrating LLM APIs across 5+ sources, reducing downstream data errors by ~35–40%. Shipped prompt engineering workflows and Power BI dashboards saving ~4–5 hrs/week. Wrote pytest coverage and structured logging across production pipelines.
Generative AI & ML Intern · Linux World Pvt. Ltd. — Jaipur · Jul–Aug 2024
Engineered ML data pipelines and applied GenAI techniques across two projects — an NLP chatbot on LLM APIs and a regression-based price predictor. Preprocessing and feature engineering in Pandas/NumPy.
📫 tarunsinghchauhan088@gmail.com | 💼 Open to AI Engineer / GenAI Engineer / LLM Engineer roles — India, Remote & Global Sponsored

