MSc Data Science at Universitat Trier. Previously four years of backend engineering at SAP Labs India. I build ML systems end to end, from training and experimentation to production deployment.
I work across the full ML stack: training and experimentation, evaluation, and production deployment. I am finishing my MSc Data Science at Universitat Trier (grade 1.5) after four years of backend engineering at SAP Labs India. My current focus is LLM efficiency and evaluation: KV cache compression, calibration and selective abstention, chain-of-thought faithfulness, and neuro-symbolic methods. I care about reproducible results, so I lean on multi-seed evaluation, ablations, and honest reporting of what did not work.
- MSc Data Science at Universitat Trier, 85 credits completed, grade 1.5, expected completion September 2026.
- MSc thesis on KV cache compression for LLM inference under Prof. Dr. Volker Schulz, submission June to July 2026. Early results show substantially lower reconstruction error than current published methods at the same compression ratio. Paper in preparation, technique details withheld until publication.
- Student Assistant for Numerical Optimization under Prof. Schulz, summer 2026.
- Independent research on LLM calibration, chain-of-thought faithfulness, and neuro-symbolic narrative modeling (see Featured Projects below).
- Learning German (currently B1.1).
Does punishing wrong answers more heavily than rewarding correct ones make a model better calibrated and more willing to abstain when it is uncertain? I built a complete GRPO training and calibration-evaluation pipeline on Qwen 2.5 1.5B and 3B, then swept seven reward conditions across six reasoning benchmarks. The empirical answer was no, and the explanation became the contribution. Finding O: under GRPO with within-group z-score advantage normalization, the entire reward design collapses to a single abstention-threshold dial, so per-question calibration cannot improve. Confirmed at 3B scale with a normalization ablation and a cross-task transfer test.
Does a reflect-and-revise step make chain-of-thought reasoning more faithful, or does it mainly make rationalizations more convincing? This study extends the counterfactual simulatability setup from Turpin et al. 2023 to reflective chain of thought, across 13 BIG-Bench Hard categories and 650 instances on Gemini 2.5 Flash Lite and LLaMA 3.2 3B. Reflective CoT cut average bias susceptibility by about 7 percentage points versus standard CoT, but did not remove systematic unfaithfulness.
LENS, my system for SemEval-2026 Task 4 (Narrative Story Similarity). An LLM decomposes each story into themes, actions, and outcomes, a heterogeneous GNN learns a graph embedding from that structure, and an additive fusion combines it with a text embedding so the final similarity score stays interpretable per modality. 72.5% pairwise accuracy and 95% cross-lingual precision across five languages.
Two-Tower MLP + FAISS retrieval on the Amazon Electronics Reviews 2023 dataset. NDCG@10 = 0.334 (+9.5% over the collaborative filtering baseline), sub-100ms inference latency on CPU, deployed on Azure Container Apps.
LlamaIndex + Google Gemini 2.5 Flash RAG over my resume and personal documents. FastAPI backend, containerized with Docker, deployed on Google Cloud Run. Live at rahul-krishnan.is-a.dev.
More projects: Sentiment Analysis with BERT (fine-tuned BERT on 10k+ YouTube comments, F1 = 0.82, served via Flask + Docker with a GitHub Actions CI/CD pipeline) and Risk-Averse Optimization case studies (CVaR-based nonlinear optimization, robust SVM, and wine-fermentation MPC, where the CVaR MPC controller held constraint violations at 0% against a baseline that violated on every trajectory).
Languages
ML & Deep Learning
LLM & RL
MLOps & Deployment
Data & Databases
- Portfolio: rahul-krishnan.is-a.dev
- LinkedIn: linkedin.com/in/rahulk98
- Email: rahulkrishnan1105@gmail.com




