每日 Harness 开源 · Source
返回本期 · Back to 2026-07-09

论文 · Papers2026-07-09 · Thursday, July 9, 2026

FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference

arxiv.org原文 ↗

上下文工程测试时计算系统·基础设施
FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference
FreqDepthKV 针对长上下文推理中的 KV cache 内存与带宽瓶颈,把相邻层 KV states 分解成共享低频 depth components 和稀疏高频 residuals。它用轻量 online probe 判断 attention heads 应使用 shared-depth、residual-depth 还是 exact cache mode,从而不用 retraining 就能随 prompt 结构调整压缩策略。在 32k-token prefill window 下,摘要报告 58.3 Exact Match、63.0 F1、32.5 ROUGE-L、48.1 pass@1,并把 decoding throughput 提到 70.4 tokens/s、TTFT 降到 2.06s、peak KV memory 降到 6.2GB,等效压缩比 3.9x。
浏览

评论 · Comments