Skip to content

perf(qwen3-omni): project only final prefill logits - #446

Merged
CjhHa1 merged 5 commits into
Tencent-Hunyuan:mainfrom
CjhHa1:perf/selective-trainside-logits
Sep 17, 2026
Merged

CjhHa1 merged 5 commits into
Tencent-Hunyuan:mainfrom
CjhHa1:perf/selective-trainside-logits

Conversation

@CjhHa1

@CjhHa1 CjhHa1 commented Sep 13, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Avoid materializing prompt-wide vocabulary logits in the Qwen3-Omni trainside prefill when generation samples only the final row.

The current Transformers 5.6 Qwen3-Omni Thinker forward does not expose logits_to_keep. A scoped helper selects hidden_states[:, -1:, :] at the lm_head boundary without replacing the module or changing its state dict. The hook exists only around the first prefill forward; decode and replay remain unchanged.

Related Issue

N/A

Test Plan

  • SKIP=no-commit-to-branch pre-commit run --all-files --show-diff-on-failure — pass
  • focused construction check: scoped LM-head selection returns [B, 1, V], matches the full projection's final row, and removes its hook after the forward
  • H20 LM-head A/B microbenchmark (B=4, L=512, H=2048, V=152064, bf16, task selective-logits-head-bench-0913): peak allocation 0.580 GiB -> 1.160 MiB; median 9.09 ms -> 0.202 ms; selected logits exact
  • full Qwen3-Omni Thinker first-prefill A/B on 8x H20, PyTorch 2.10, Transformers 5.6, bf16, Accelerate balanced device map, input [4, 512] (task pr446-full-prefill-ab-8h20-0913-v3): max per-GPU forward allocation 0.596 -> 0.307 GiB (-0.289 GiB, -48%); median 259.14 -> 254.48 ms (+1.8%); selected logits close and sampled KV cache exact
  • production FSDP2 full-shard first-prefill A/B on 8x H20, one rank/GPU, production LoRA/wrap settings, input [4, 512] (task pr446-fsdp-qwen-8h20-0914): max-rank forward peak 3.958 -> 3.958 GiB (unchanged because FSDP all-gather dominates); median max-rank latency 363.03 -> 344.13 ms (+5.5%); selected logits close and sampled KV cache exact on every rank
  • full token-generation/training loop: not run; the changed first-prefill forward is isolated above with the real checkpoint and production FSDP wrapping

Compatibility / Risk

No config, checkpoint, state-dict, or public model API changes. Prompts are left-padded before generation, so the final hidden row is each sample's last real token. The forward pre-hook is removed in finally; subsequent decode steps and replay do not carry it.

Under production FSDP the optimization improves prefill latency, but does not reduce the observed peak because the larger FSDP layer all-gather determines that peak.

Reviewer Notes

The native Qwen logits_to_keep changes live in #442. This PR covers Qwen3-Omni because its current forward lacks that argument.

HunyuanImage3 was benchmarked and removed from this PR: under its production FSDP topology, peak changed only 9.474 -> 9.471 GiB and latency regressed 2672.10 -> 2707.63 ms, so the scoped projection did not justify its added complexity.

Checklist

  • I reviewed the changed code and removed unrelated/generated artifacts.
  • I updated tests, docs, and configs where needed, or explained why not.

@github-actions github-actions Bot added the need review Ready and waiting for review label Sep 13, 2026
@CjhHa1
CjhHa1 force-pushed the perf/selective-trainside-logits branch from f4f275d to 2bd764c Compare September 13, 2026 04:22
@CjhHa1 CjhHa1 changed the title perf(ar): project only sampled logits during trainside generation perf(ar): project selected prefill logits for HI3 and Qwen3-Omni Sep 13, 2026
@CjhHa1
CjhHa1 force-pushed the perf/selective-trainside-logits branch from 2bd764c to 825b6da Compare September 14, 2026 01:13
@CjhHa1 CjhHa1 changed the title perf(ar): project selected prefill logits for HI3 and Qwen3-Omni perf(qwen3-omni): project only final prefill logits Sep 14, 2026
@github-actions github-actions Bot added approved Approved by reviewer and removed need review Ready and waiting for review labels Sep 16, 2026
@CjhHa1
CjhHa1 merged commit 852f151 into Tencent-Hunyuan:main Sep 17, 2026
5 checks passed
@CjhHa1
CjhHa1 deleted the perf/selective-trainside-logits branch September 17, 2026 02:00
@github-actions github-actions Bot removed the approved Approved by reviewer label Sep 17, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants