Grok 4.6 high benchmark results - #28
Conversation
|
this looks much more expensive than Grok 4.5. no idea why |
The cost increase comes from two effects:
For example, on Task 01, both models ran for approximately 30 minutes, but Grok 4.5 used 47 steps and 1.52M cumulative input tokens, while Grok 4.6 used 93 steps and 6.58M input tokens. Grok 4.6 performed more exploration, testing, and tool calls; each subsequent call carried an increasingly long conversation and tool-output history, raising cumulative input by approximately 4.3× on that task. |
Summary
Adds the complete Grok 4.6 (high) ORBIT-Q TensorCircuit benchmark archive and publication-ready result figures.
grok-4.6, reasoning efforthighgpt-5.6-sol, reasoning efforthigh1.8.0.dev20260726sha256:b059c5fa7f75702f9afbf94ec7866e102ac32afd59d25634ec0aca0fd56e283319fe27b83eaf668b3df32d1a68902b08cbe28585f189290769018eb16d9278950201238ec2983907e2891f5319f5fff2d00844d5All 12 selected outcomes are first attempts under the repaired xAI compatibility layer. No integer-tool-schema or reasoning-replay compatibility failures recurred. Task 01 reached the 1,800-second agent limit; Tasks 05 and 08 passed functional checks but failed the independent source audit. Task 08 is also negative under the campaign-wide human adjudication.
Results
The archive includes selected solutions, rewards, solver trajectories, verifier/audit details, runtime and token accounting, configs, task/image provenance, PNG/PDF/SVG figures, and reproduction/verification tools.
Resource accounting
Harbor's raw
cost_usdfield is preserved as provider-reported zero/null; the separate reconstructed cost is derived from the archived token classes.Validation
Verified: 12 valid outcomes, fixed task hashes, Grok 4.6/high solver, GPT-5.6 Sol/high audit, pinned TensorCircuit 1.8 image, solution hashes, and valid figures.