I've made lot of work to significantly improve Kvarn KV Cache performance on my RTX 3090 for Qwen dense and MoE models, especially for concurrent requests without losing quality. Right now, the performance is on par with Q8 KV and Kvarn KV is faster on large context than Q8 KV. I hope, you will find my contribution useful. Since I've done it for myself, currently all docs and comments on my changes are in Russian.
https://github.com/valujin/beellama-kvarn
I've made lot of work to significantly improve Kvarn KV Cache performance on my RTX 3090 for Qwen dense and MoE models, especially for concurrent requests without losing quality. Right now, the performance is on par with Q8 KV and Kvarn KV is faster on large context than Q8 KV. I hope, you will find my contribution useful. Since I've done it for myself, currently all docs and comments on my changes are in Russian.
https://github.com/valujin/beellama-kvarn