Skip to content

Gemma 4 [6/6]: validation — HF logit/argmax parity + perf + memory on Orin #94

Description

@ai-hpc

Part of #88. Depends on [4/6], [5/6].

Scope

  • Greedy-decode logit/argmax parity vs HF transformers Gemma 4 E2B on a fixed prompt set (top-1 token match; bounded logit drift).
  • Coherence check (interactive generation).
  • tok/s (prefill + decode) and peak memory (model + KV + PLE table) within the 8 GB budget shared with voice.

Acceptance

  • Top-1 parity vs HF on the prompt set; documented perf/memory; ready to port the math to TensorRT-Edge-LLM#72.

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Fields

    No fields configured for issues without a type.

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions