I work across the whole deployment loop: technical discovery, POCs and evals, model training, agent runtimes, GPU infrastructure, realtime interfaces, and the production traces that make the next iteration better.
That range comes from operating across engineering, solutions, GTM, and production support—from Tesla and Labelbox through First Round and UW, into the systems I build today.
customer problem → prototype → held-out eval → model / agent → governed runtime → production evidence
↑ │
└──────────────── improve ─────────────────────┘
Karti-Small-RSI-3B targets reliable tool calling and offline agent workflows; Karti-Small-VL-4B adds vision grounding and abstention for unreadable identifiers. The 3B model card publishes its pinned foundation, BF16 LoRA SFT recipe, evaluation contract, and human-gated improvement loop while weights and training rows stay private; the Apache-2.0 4.66B BF16 and vision-preserving NVFP4 releases are public.
PyTorch TRL LoRA / SFT Prime Intellect Verifiers Hugging Face GH200 DGX Spark / GB10 NVFP4 tool calling
Compute and evaluation foundations paired with private agent worlds: governed unified-memory AI nodes, verifier-driven environments, deterministic simulation, presence, and workspace systems.
Rust Python TypeScript React three.js MCP DGX Spark / GB10
A personalized trading harness wrapped around one human trader: realtime market interfaces, Python/Rust execution, risk controls, a self-hosted data plane, MCP tools, and voice agents sharing one live operating picture.
Python Rust React PostgreSQL WebSockets LiveKit MCP
An ambient voice-AI pair programmer that watches a shared engineering room and surfaces duplicated work, merge collisions, and missed handoffs. Built and shipped with a first-time hackathon team at the 2026 AI Engineer World’s Fair.
LiveKit Gemma Modular MAX MongoDB DigitalOcean
| Project | What it proves | Core stack |
|---|---|---|
| comma-controls-challenge | A causal realtime lateral controller built by measuring the simulator, fitting feedforward, and using the future plan as a zero-lag smoother. | Python · ONNX · controls |
| opencode-filter | A fail-closed I/O boundary that detects and replaces secrets before agent traffic reaches a model. | TypeScript · HMAC · entropy detection |
| karti-code | An agentic coding harness that composes specialized agents, lifecycle hooks, and private infrastructure tools. | TypeScript · OpenCode · MCP |
| search-mcp-server | Local-first web and code search for agents through SearXNG and Grep.app. | Go · SearXNG · Docker |
| gitea-mcp-server | Repository operations exposed as structured tools for agent workflows. | Go · Gitea · MCP |
| caddy-mcp-server | Safe reverse-proxy inspection and operations through an agent-facing control plane. | Go · Caddy · MCP |
| openclaw-mattermost-extension | Agent orchestration connected to the place engineering teams already coordinate. | TypeScript · Mattermost · OpenClaw |
| AutoMagically | Automated single-GPU research loops built from Karpathy's Auto Research work. | Python · PyTorch · experiment loops |
models & evals PyTorch · TRL · LoRA/SFT · verifier-driven RL · dataset curation
agent systems tool calling · MCP · LiveKit · OpenCode · Claude Code · Codex
runtime vLLM · SGLang · llama.cpp · CUDA · NVFP4 · quantization · offline inference
systems Rust · Python · TypeScript · Go · React · WebSockets · PostgreSQL
infrastructure Docker · Kubernetes · Caddy · Tailscale · Proxmox · GitHub Actions
field work discovery · solution architecture · POCs · integrations · production support
- Making local agents dependable enough to operate tools, not just narrate intentions.
- Building evaluation environments that produce useful evidence instead of leaderboard theater.
- Turning GPU nodes into governed, agent-operable infrastructure.
- Closing the loop between realtime human work, agent behavior, and the next training run.
Build the system. Measure the behavior. Improve the loop.
karti.ai ·
resume ·
email