Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills
arxiv.org原文 ↗
ACES 把 skill、bundle、plugin 当作可执行能力包,固定模型、工作区、harness 和 scorer 做有无 skill 的配对试验,并以 ATIF 统一轨迹。145 个 skill 的扫描指标与 LLM judge 仅有 rho=0.14;947 个配对案例平均 composite Skill Lift 0.2134,72.8% 案例为正,说明静态文档门无法代替运行时测量。
–浏览
评论 · Comments