Skip to content

Disable ORT graph optimization: session load alone jetsammed the app - #91

Merged
sanylax0 merged 1 commit into
mainfrom
claude/repo-review-improvements-yvk3t7-ort-load-explosion
Jul 18, 2026
Merged

sanylax0 merged 1 commit into
mainfrom
claude/repo-review-improvements-yvk3t7-ort-load-explosion

Conversation

@sanylax2

Copy link
Copy Markdown
Collaborator

What

The heartbeat run pinned the killer definitively: plain playback is flat (~3323 MB headroom, no leak), the separation gate passed with 3.3 GB free, and the app died between ort session load begin and ort session ready — ORTSession creation alone consumed over 3.3 GB.

Why

The model ships fp16 weights behind per-weight Cast nodes. ORT's default load-time graph optimization constant-folds those casts, materializing an fp32 copy of every weight while the protobuf buffer and the fp16 originals are still resident — several full copies of an 85M-param transformer at once.

How

setGraphOptimizationLevel(.none) in sharedSession() (enum/API verified against onnxruntime v1.20.0's ObjC headers). Per-window inference gets somewhat slower without fused ops — fine for an offline cache-once job (same trade as choosing the CPU EP). If load now fits but windows are unacceptably slow, the follow-up is an offline-optimized .ort model, which bakes the optimizations in without the load-time explosion.

Testing

On device: play an un-separated track; after the 60 s hold, the console should now progress ort session load begin → ort session ready → window 0, 10, … with the heartbeat staying healthy. Paste the trace if it dies anywhere new.

🤖 Generated with Claude Code

https://claude.ai/code/session_013TJoWkqg8bzkGdzxWhWjWP


Generated by Claude Code

The playback heartbeat pinned it: headroom is flat (~3323 MB) through
plain playback, the separation gate passed with 3.3 GB free, and the
app died between 'ort session load begin' and 'ort session ready' —
session creation alone consumed over 3.3 GB. The fp16 model keeps its
weights behind per-weight Cast nodes, and load-time constant folding
materializes an fp32 copy of every weight while the protobuf and fp16
originals are still resident. Skip graph optimization entirely: slower
per window (fine for an offline cache-once job), but the load peak
actually fits.
@sanylax0
sanylax0 merged commit bd3481a into main Jul 18, 2026
1 of 4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants