每日 Harness 开源 · Source
返回本期 · Back to 2026-06-20

论文 · Papers2026-06-20 · Saturday, June 20, 2026

Context-Aware RL for Agentic and Multimodal LLMs

arxiv.org原文 ↗

Context-Aware RL for Agentic and Multimodal LLMs
这篇提出 ContextRL,用 context-aware reinforcement learning 改善长程推理和多模态表现。摘要里的失败例子很具体:答案可能取决于工具 trace 里的一行,或图像里一个细微细节;方法用 indirect auxiliary objective 训练模型识别这种决定性证据。对 agentic LLM 来说,这类训练目标直接对应工具使用、证据查找和多轮任务中的注意力分配问题。
浏览

评论 · Comments