FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference
arxiv.org原文 ↗
FreqDepthKV 针对长上下文推理中的 KV cache 内存与带宽瓶颈,把相邻层 KV states 分解成共享低频 depth components 和稀疏高频 residuals。它用轻量 online probe 判断 attention heads 应使用 shared-depth、residual-depth 还是 exact cache mode,从而不用 retraining 就能随 prompt 结构调整压缩策略。在 32k-token prefill window 下,摘要报告 58.3 Exact Match、63.0 F1、32.5 ROUGE-L、48.1 pass@1,并把 decoding throughput 提到 70.4 tokens/s、TTFT 降到 2.06s、peak KV memory 降到 6.2GB,等效压缩比 3.9x。
–浏览
评论 · Comments