每日 Harness 开源 · Source
返回本期 · Back to 2026-07-23

论文 · Papers2026-07-23 · Thursday, July 23, 2026

RECON: Benchmarking Agent Memory for Compositional Reasoning over Long Contexts

arxiv.org原文 ↗

RECON: Benchmarking Agent Memory for Compositional Reasoning over Long Contexts
RECON 不是只测试记住某句话,而是要求智能体跨长对话组合多个分散事实完成推理。基准包含 1,000 个任务、最长 100 万 token 的交互历史,并区分检索、整合和答案生成环节。该设计能暴露“检索到了但拼不起来”的记忆失败,比单轮 needle-in-a-haystack 更接近持续运行代理的真实负载。
浏览

评论 · Comments