Any chance you can publish a standard trace format (prompt versions, tool calls, token/time, and what the human grader saw) so others can replicate results?
Also, I would be curious to see a separate score for "safety" or "hallucination cost" when the agent is generating something that looks like an engineering recommendation.
Related: https://www.agentixlabs.com/ has a couple posts on agent eval harnesses and what to log if you want to make benchmarks reproducible.
Any chance you can publish a standard trace format (prompt versions, tool calls, token/time, and what the human grader saw) so others can replicate results?
Also, I would be curious to see a separate score for "safety" or "hallucination cost" when the agent is generating something that looks like an engineering recommendation.
Related: https://www.agentixlabs.com/ has a couple posts on agent eval harnesses and what to log if you want to make benchmarks reproducible.