I build the tooling layer between LLM APIs and production - cost estimation, context grounding, output stability.
- Token Estimator — pre-deployment cost planning CLI
- ContextEval — RAG grounded strictly to your corpus
- TraceBench — output stability and reasoning variance analyzer