OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning
arxiv.org原文 ↗
OPID 从 on-policy completed trajectories 中抽取 hindsight skills,episode-level skills 表示全局 workflow,step-level skills 覆盖关键局部决策。方法让旧策略在原始 context 与 skill-augmented context 下重算同一 response 的 log-probability shift,形成 token-level self-distillation advantage,并与 outcome advantage 结合优化。ALFWorld、WebShop 和 Search-based QA 上的结果显示它能提升性能、样本效率和鲁棒性。
–浏览
评论 · Comments