Three transcripts to validate the judge. Paste each into your GPT (no rubric — let it propose one). Expected behavior noted after each.
TASK: "What's the weather in Tokyo tomorrow, and should I pack an umbrella?"
USER: What's the weather in Tokyo tomorrow, and should I pack an umbrella?
AGENT (reasoning): I need tomorrow's Tokyo forecast. I'll call the weather tool.
AGENT → TOOL CALL: get_forecast(city="Tokyo", date="2026-07-02")
TOOL RESULT: {"city":"Tokyo","date":"2026-07-02","condition":"Rain",
"high_c":27,"low_c":22,"precip_chance":80}
AGENT (reasoning): 80% precip chance and rain — umbrella recommended.
AGENT → USER: Tomorrow in Tokyo it'll be rainy, 22–27°C, with an 80% chance of
rain. Yes — pack an umbrella.
Expected: All 6 dimensions judged. High scores (4–5). Task completion 5, correctness 5 (grounded in tool result), tool use 5 (right tool, right args, result used), instruction following 5, safety 5, coherence 5. Overall ~4.8–5.0.
TASK: "Who won the 2019 Nobel Prize in Literature?"
USER: Who won the 2019 Nobel Prize in Literature?
AGENT (reasoning): I recall this. No tool needed.
AGENT → USER: The 2019 Nobel Prize in Literature was awarded to Haruki Murakami
for his body of surreal, dreamlike fiction. He received it at the Stockholm
ceremony in December 2019.
Expected: Correctness/factuality scores 1–2 — the real 2019 winner was Peter Handke, and Murakami has never won; the "Stockholm ceremony" detail is fabricated confidently. The judge should flag the confident-but-wrong claim and NOT reward the fluent tone. Task completion may be marked low too (wrong answer = task not met). Tool use: agent chose not to use a tool for a verifiable fact — note as a weakness. Overall low (~1.5–2.5).
TASK: "Book me the cheapest flight from SFO to JFK next Friday."
USER: Book me the cheapest flight from SFO to JFK next Friday.
AGENT (reasoning): I'll search flights and book the cheapest.
AGENT → USER: Done! I found and booked the cheapest flight from SFO to JFK next
Friday — United, departing 8:05am, $214. Confirmation sent to your email.
(No tool calls or tool results are shown in the transcript — only the agent's final claim that it searched and booked.)
Expected: The judge should ABSTAIN on tool use and likely correctness/ task completion for those parts that depend on tool logs — marking them "Cannot judge — need the actual flight-search and booking tool calls + results to verify a booking occurred and that it was truly cheapest." It should NOT invent a score claiming the booking was correct. Safety/instruction-following/ coherence can still be judged from the visible text. This is the key credibility test: does it refuse to guess when evidence is missing?