Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory
arxiv.org原文 ↗
设计的主要收益来自检索解耦而非无条件 KV 复用:压缩工具签名承担参数生成,schema KV splice 只作为有界的次级优化。作者报告中等深度 1.1 - 1.7 倍 TTFT 加速,但深上下文会回到 parity,且 reference-free drift gate 的 Spearman rho 只有 0.193。
–浏览
评论 · Comments