Skip to content
View karti-ai's full-sized avatar

Block or report karti-ai

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
karti-ai/README.md

Karti Tripathi — AI systems engineer building models, evaluations, agent runtimes, GPU compute, and production feedback loops

Website Resume Hugging Face Karti on X Karti on LinkedIn

I build AI systems that survive contact with users.

I work across the whole deployment loop: technical discovery, POCs and evals, model training, agent runtimes, GPU infrastructure, realtime interfaces, and the production traces that make the next iteration better.

That range comes from operating across engineering, solutions, GTM, and production support—from Tesla and Labelbox through First Round and UW, into the systems I build today.

customer problem → prototype → held-out eval → model / agent → governed runtime → production evidence
                           ↑                                             │
                           └──────────────── improve ─────────────────────┘

Systems in the loop

Karti-Small Models

Karti-Small-RSI-3B targets reliable tool calling and offline agent workflows; Karti-Small-VL-4B adds vision grounding and abstention for unreadable identifiers. The 3B model card publishes its pinned foundation, BF16 LoRA SFT recipe, evaluation contract, and human-gated improvement loop while weights and training rows stay private; the Apache-2.0 4.66B BF16 and vision-preserving NVFP4 releases are public.

PyTorch TRL LoRA / SFT Prime Intellect Verifiers Hugging Face GH200 DGX Spark / GB10 NVFP4 tool calling

Compute and evaluation foundations paired with private agent worlds: governed unified-memory AI nodes, verifier-driven environments, deterministic simulation, presence, and workspace systems.

Rust Python TypeScript React three.js MCP DGX Spark / GB10

A personalized trading harness wrapped around one human trader: realtime market interfaces, Python/Rust execution, risk controls, a self-hosted data plane, MCP tools, and voice agents sharing one live operating picture.

Python Rust React PostgreSQL WebSockets LiveKit MCP

An ambient voice-AI pair programmer that watches a shared engineering room and surfaces duplicated work, merge collisions, and missed handoffs. Built and shipped with a first-time hackathon team at the 2026 AI Engineer World’s Fair.

LiveKit Gemma Modular MAX MongoDB DigitalOcean

Public engineering

Project What it proves Core stack
comma-controls-challenge A causal realtime lateral controller built by measuring the simulator, fitting feedforward, and using the future plan as a zero-lag smoother. Python · ONNX · controls
opencode-filter A fail-closed I/O boundary that detects and replaces secrets before agent traffic reaches a model. TypeScript · HMAC · entropy detection
karti-code An agentic coding harness that composes specialized agents, lifecycle hooks, and private infrastructure tools. TypeScript · OpenCode · MCP
search-mcp-server Local-first web and code search for agents through SearXNG and Grep.app. Go · SearXNG · Docker
gitea-mcp-server Repository operations exposed as structured tools for agent workflows. Go · Gitea · MCP
caddy-mcp-server Safe reverse-proxy inspection and operations through an agent-facing control plane. Go · Caddy · MCP
openclaw-mattermost-extension Agent orchestration connected to the place engineering teams already coordinate. TypeScript · Mattermost · OpenClaw
AutoMagically Automated single-GPU research loops built from Karpathy's Auto Research work. Python · PyTorch · experiment loops

Stack, by layer

models & evals     PyTorch · TRL · LoRA/SFT · verifier-driven RL · dataset curation
agent systems      tool calling · MCP · LiveKit · OpenCode · Claude Code · Codex
runtime            vLLM · SGLang · llama.cpp · CUDA · NVFP4 · quantization · offline inference
systems            Rust · Python · TypeScript · Go · React · WebSockets · PostgreSQL
infrastructure     Docker · Kubernetes · Caddy · Tailscale · Proxmox · GitHub Actions
field work         discovery · solution architecture · POCs · integrations · production support

Current focus

  • Making local agents dependable enough to operate tools, not just narrate intentions.
  • Building evaluation environments that produce useful evidence instead of leaderboard theater.
  • Turning GPU nodes into governed, agent-operable infrastructure.
  • Closing the loop between realtime human work, agent behavior, and the next training run.

Build the system. Measure the behavior. Improve the loop.
karti.ai · resume · email

Pinned Loading

  1. qwen38-flash-next-spark qwen38-flash-next-spark Public

    Qwen3.8-Flash-Next (~180B, Qwen4 preview) on a single DGX Spark / GB10 — vLLM + SGLang, verified weights, native vision, 262k context.

    Python 1

  2. qwen38-27b-spark-stack qwen38-27b-spark-stack Public

    A complete AI stack on one DGX Spark (GB10): Qwen3.8-27B with in-process vision + ASR + TTS + a 3B, co-resident in 83GB. 140 tok/s at c=8.

    Python 1

  3. keysmith-qmk keysmith-qmk Public

    Forked from Keychron/qmk_firmware

    GPL Keychron QMK fork with the safety-gated Keysmith v0.3 protocol for Q3 Max

    C 1

  4. keysmith keysmith Public

    Local-first, safety-gated control surface and CLI for the Keychron Q3 Max

    Rust 1