每日 Harness 开源 · Source
返回本期 · Back to 2026-08-27

论文 · Papers2026-08-27 · Thursday, August 27, 2026

Joint Optimization of Tool Creation and Use for Large Language Model Agents

arxiv.org原文 ↗

Joint Optimization of Tool Creation and Use for Large Language Model Agents
SMITH 在单一策略中交替执行“从示例写工具”的 build rollout 与“调用池化工具”的 use rollout,并为 schema、代码、结果设置独立奖励。4B Qwen3 在 13 个精确验证程序推理任务上达 79.8 宏平均准确率,超过未训练的 30B-A3B 工具编写器;TabMWP-Hard 40.4、OOD GQA 42.6,比同骨干推理时基线高 7.6 点。工具质量还会外溢提升 LFM-2.5-350M 和 Qwen3-30B-A3B。
–浏览

评论 · Comments