Prerequisites
Feature Description
Please implement adaptive-kv-streaming
Motivation
I can run higher quants for both model and KV and have a much higher context with minimal drop in performance.
Possible Implementation
Add a flag to enable or disable it.
Prerequisites
Feature Description
Please implement adaptive-kv-streaming
Motivation
I can run higher quants for both model and KV and have a much higher context with minimal drop in performance.
Possible Implementation
Add a flag to enable or disable it.
I checked the initial repo and the article and I failed to find the most important thing there: KLD or some other measurements of how this affects output. If you are already a user of this fork, could you please benchmark that, or create an issue asking for that?
The thing is, if output quality does not match the full resident even roughly, then there's barely any point in the feature as a whole.