Building observable, verifiable AI systems — from agent runtimes to scientific workflows.
专注 Agent 工程、开发者工具,以及可观测、可验证的 AI 工作流。
Selected Work · Open Source · Stack · Email
|
01 · ORCHESTRATE Reliable agent systems with explicit lifecycle, routing, and human-control boundaries. |
02 · OBSERVE Safe MCP tooling that turns runtime signals into compact, privacy-aware context. |
03 · VERIFY Evidence-backed automation and reproducible AI workflows that fail clearly. |
写得越少,bug 越少。自动化也应该可观察、可验证、可停止。
|
A Python agent runtime powered by TypeSafe Jev: guarded tool execution, isolated Docker sandboxes, and side-by-side LLM comparisons. Python Agent Runtime Docker
|
An observable research agent for reproducible single-cell RNA-seq analysis. LangGraph FastAPI React
|
|
Maker-checker goal loops for bounded, evidence-backed coding-agent work. Agent Skills Verification
|
A local MCP runtime for asynchronous delegation to external CLI coding agents. MCP Node.js CLI Agents
|
|
A read-only, privacy-aware MCP server for querying and summarizing Jaeger traces. Python MCP Observability
|
A full-stack FastAPI + LangGraph starter with React, SSE, and persistent checkpoints. FastAPI LangGraph SSE
|
More experiments → RAG, MLX, biomedical NLP, and developer-tooling explorations
|
An agent development environment for working with fleets of parallel coding agents. TypeScript Electron MIT
|
A benchmark for evaluating LLM × harness performance. Python Agent Evaluation Apache-2.0
|
|
A cross-platform desktop assistant for switching API providers across Claude Code, Codex, and other AI coding agents. Rust Tauri MIT
|
- AI & agents:
LangGraph·LangSmith·PyTorch·NumPy·MLX·MCP - Backend & product:
FastAPI·Django·React·PostgreSQL - Languages & web:
Python·TypeScript·C·Shell·HTML5·CSS3 - Systems & workflow:
Linux·macOS·Docker·Git·SONiC·Cursor
If you're working on agent reliability, MCP tooling, or applied AI, feel free to get in touch.

