State-aware hedged requests for serverless GPU inference — return the first valid result and cancel the losers with an audited receipt.
python modal inference cold-start speculative-execution mlops serverless-gpu runpod hedged-requests cerebrium
-
Updated
Jul 16, 2026 - Python