Skip to content

BF16 / FP16 precision support in NKI kernels #50

Description

@scttfrdmn

All current NKI kernels are FP32. Trainium natively supports BF16 and FP16 with meaningful throughput advantages (~2× on Tensor Engine for BF16).

Deliverable: dtype-parametric versions of _complex_gemm_kernel, _complex_linear_kernel, _complex_mul_kernel, butterfly_stage_kernel. Tests with tolerance appropriate to the reduced precision.

Benchmark expected: 1.5-2× speedup on large GEMM and STFT paths.

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestnkiNKI kernel development (Trainium hardware)

    Type

    No type

    Projects

    No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions