Skip to content

DeepSeek V4 Pro high benchmark results - #27

Open
QingyunQian wants to merge 1 commit into
sxzgroup:v2from
QingyunQian:codex/deepseek-v4-pro-high-tc18-results
Open

DeepSeek V4 Pro high benchmark results #27
QingyunQian wants to merge 1 commit into
sxzgroup:v2from
QingyunQian:codex/deepseek-v4-pro-high-tc18-results

Conversation

@QingyunQian

Copy link
Copy Markdown

Summary

Adds the complete DeepSeek V4 Pro (high) ORBIT-Q TensorCircuit benchmark archive and publication-ready result figures.

  • Solver: deepseek-v4-pro, reasoning effort high
  • Auditor: gpt-5.6-sol, reasoning effort high
  • Result: 5/12 passed — Tasks 02, 04, 06, 07, 10
  • TensorCircuit: exactly 1.8.0.dev20260726
  • Immutable image ID: sha256:b059c5fa7f75702f9afbf94ec7866e102ac32afd59d25634ec0aca0fd56e2833
  • Resources: 6 CPU, 10240 MiB reservation, 16384 MiB limit
  • Frozen task aggregate SHA-256: 19fe27b83eaf668b3df32d1a68902b08cbe28585f189290769018eb16d927895
  • Benchmark base commit: 0201238ec2983907e2891f5319f5fff2d00844d5

The earlier DeepSeek V4 Pro campaign run with TensorCircuit 1.7 is explicitly excluded; this PR contains only the clean TC 1.8 campaign selected for reporting.

Results

DeepSeek V4 Pro outcomes

DeepSeek V4 Pro agent-side resource use

The archive includes selected solutions, rewards, solver trajectories, verifier/audit details, runtime and token accounting, configs, task/image provenance, PNG/PDF/SVG figures, and reproduction/verification tools.

Validation

python3 results/deepseek-v4-pro-high/tools/verify_archive.py

Verified: 12 valid outcomes, fixed task hashes, DeepSeek V4 Pro/high solver, GPT-5.6 Sol/high audit, pinned TC 1.8 image, and valid figures.

@QingyunQian QingyunQian changed the title DeepSeek V4 Pro high benchmark results (TensorCircuit 1.8) DeepSeek V4 Pro high benchmark results Aug 13, 2026
@QingyunQian
QingyunQian marked this pull request as ready for review August 13, 2026 15:17
@refraction-ray

Copy link
Copy Markdown
Member

thanks for the nice contribution, though the model looks so-so

@refraction-ray

refraction-ray commented Aug 14, 2026

Copy link
Copy Markdown
Member

seems that luna is the option wayyy better than deepseek models (both pro and flash) in terms of time cost, money cost, intelligence and capability. very inspiring results

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants