Skip to content
View karanparekh14's full-sized avatar

Block or report karanparekh14

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
karanparekh14/README.md

Hi, I'm Karan Parekh

AI Developer and Data Analyst | MS Advanced Data Analytics (AI) @ UNT (2026) | MS Data Science @ LJMU (2025)

I build machine learning pipelines, production AI platforms, and autonomous agent systems. About 5.5 years of industry experience across technical delivery, analytics, automation and operations, now building AI-powered products and applying ML to real business problems.


Publications

Year Work Where
2026 When the Agent Speaks for the Company: A Twelve-Week Experience Report on Failure Boundaries and Control Design in a Multi-Party LLM Agent Zenodo, doi:10.5281/zenodo.22704694 · arXiv submitted
2026 When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document Contamination Zenodo, doi:10.5281/zenodo.21939087 · arXiv:2609.09696

When the Agent Speaks for the Company is a single-author field study of a large language model agent that ran for twelve weeks inside a working business, speaking to staff, partner staff and outside counterparties. It codes 45 logged incidents by causal mechanism and impact, measures a confidentiality leak rate against the full message record, and quantifies how much the incident log itself undercounts. The coded incident record is released as a CSV supplement.

When Auditors Fabricate tests whether a frontier model can find deliberately planted errors in academic papers. It reports where that ability collapses, and shows that the failure mode at scale is confident fabrication rather than abstention.

Both are open access under CC BY 4.0. The links above are concept DOIs and always resolve to the newest version.

ORCID: 0009-0000-1681-3274


What I Build

Production AI Systems

  • Built and deployed SourceWithAI, a full-stack AI-powered B2B sourcing platform with conversational search, multi-LLM orchestration, and vector search (Next.js + Express + FastAPI + MongoDB + Elasticsearch)
  • Designed and operated an autonomous WhatsApp AI agent (Chotu) that ran business operations daily through to May 2026: product sourcing, presentation generation, email campaigns, and team coordination using persistent memory architecture

Machine Learning & Statistical Research

  • Classification, regression, NLP, XGBoost, SHAP interpretability
  • Hierarchical regression on 646K federal employee records that surfaced a Simpson's Paradox
  • LLM evaluation research testing Gemini's contamination detection on 150 academic PDFs
  • Field research on failure boundaries and control design in production LLM agents

Tech Stack

  • Languages: Python, JavaScript/TypeScript, SQL, R
  • ML/AI: Scikit-learn, XGBoost, SHAP, Statsmodels, Pandas, NumPy, OpenAI, Gemini, Anthropic Claude
  • Search & AI: Vector embeddings (Transformers.js), RAG, Elasticsearch, intent classification, LangChain
  • Web: Next.js, React, Express.js, FastAPI, Node.js, Socket.io
  • Databases: MongoDB, PostgreSQL, MySQL, Redis, Elasticsearch
  • Cloud & DevOps: AWS (S3), Vultr VPS, Nginx, PM2, Vercel, Docker
  • Tools: Git, Tableau, Power BI, Stripe API, Gamma API, OTAPI

Featured Projects

Project What It Does Stack
SourceWithAI Platform Full-stack AI-powered B2B sourcing marketplace with conversational search, multi-platform aggregation, and 50+ API endpoints. Built and run through May 2026. Next.js, Express, FastAPI, MongoDB, Elasticsearch, Redis, GPT-4o, Claude
Chotu: AI Operations Agent WhatsApp-connected autonomous agent. Product sourcing, presentation generation, email campaigns, persistent memory. Ran in production for twelve weeks to May 2026. OpenClaw, Claude Sonnet, Node.js, Gamma API, OTAPI, SmartLead
Customer Churn Prediction Interpretable XGBoost and SHAP classification pipeline for telecom churn. 78.2% accuracy, 0.848 AUC on the test set, with a costed retention business case. Python, XGBoost, SHAP, SMOTE
Federal Employee Satisfaction Hierarchical OLS regression on 646K survey records, final R2 of .597. Surfaced a Simpson's Paradox in telework data. Python, Statsmodels, Matplotlib
Lead Scoring Model Logistic regression lead scoring. 81.5% test accuracy, 76.3% recall. ROC-AUC 0.88 on train, VIF multicollinearity analysis. Python, Scikit-learn, Statsmodels
LLM Evaluation Research Tested Gemini's ability to detect 450 planted contaminants in 150 academic PDFs. Published: When Auditors Fabricate, also arXiv:2609.09696. Python, Gemini API, NLP

Connect


Currently seeking Data Analyst / Data Scientist / ML Engineer roles.

Pinned Loading

  1. llm-contamination-detection-eval llm-contamination-detection-eval Public

    Evaluating Gemini's ability to detect knowledge contamination in 150 academic PDFs with 450 planted contaminants. Graduate research project, UNT.

    Python

  2. federal-employee-satisfaction-analysis federal-employee-satisfaction-analysis Public

    Hierarchical regression on telework, work-life balance and engagement vs. federal employee satisfaction. 646K FEVS 2024 records, R2 = .597. Analytics capstone, UNT.

    Jupyter Notebook

  3. chotu-ai-agent chotu-ai-agent Public

    Production multi-agent operations platform on WhatsApp. 5 agents, 88 scheduled jobs, tiered model routing, fabrication guards and a written incident runbook. Running continuously since Feb 2026.

  4. CUSTOMER-CHURN-PREDICTION-IN-TELECOM-USING-MACHINE-LEARNING CUSTOMER-CHURN-PREDICTION-IN-TELECOM-USING-MACHINE-LEARNING Public

    Interpretable XGBoost and SHAP pipeline for telecom customer churn. 78.2% accuracy and 0.848 AUC on the held-out test set. LJMU MS Data Science thesis.

    Jupyter Notebook 1

  5. Lead-Score-Case-Study Lead-Score-Case-Study Public

    Logistic regression model to score and prioritize sales leads by conversion probability. 81.5% test accuracy, 76.3% recall, ROC-AUC 0.88 on train.

    Jupyter Notebook 1

  6. sourcewithai-platform sourcewithai-platform Public

    Production B2B sourcing platform. Staged multi-model search pipeline with dual-reality retrieval, Elasticsearch and vector hybrid ranking, supplier intelligence. 47 API domains, 294 handlers. Live …