Hello, I would like to ask how FLUX.2-klein-9b-kv was trained? Was it distilled from a normal base model without kv cache inference through special distillation, or was it distilled from a multi-step teacher model that inherently has kv cache inference capabilities? Do you have any articles or work examples for reference?
Hello, I would like to ask how FLUX.2-klein-9b-kv was trained? Was it distilled from a normal base model without kv cache inference through special distillation, or was it distilled from a multi-step teacher model that inherently has kv cache inference capabilities? Do you have any articles or work examples for reference?