RoPE-Aware Bit Allocation for KV-Cache Quantization
arxiv.org原文 ↗
这篇把 RoPE 下 key cache 的误差敏感性建成二维频率块的 bit allocation 问题,而不是把 cached key 当平面向量量化。Block-GTQ 在 K-only 2/3 b-dim 量化中将 per-layer MAE 降低 32-80%,并赢下 367/367 个 layer 对比;Llama-3.1-8B-Instruct 的 K2V2 设置下,NIAH 六任务均值从 70.6 升至 97.4,LongBench-EN 从 36.87 升至 53.31。服务路径也有工程结果:Qwen2.5-3B packed K3V3 在单 H800 上达到 3.24 倍 KV-cache 压缩,128K 上峰值内存从 56.31 GB 降到 19.85 GB。
–浏览
评论 · Comments