每日 Harness 开源 · Source
返回本期 · Back to 2026-08-25

论文 · Papers2026-08-25 · Tuesday, August 25, 2026

Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills

arxiv.org原文 ↗

Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills
ACES 把 skill、bundle、plugin 当作可执行能力包,固定模型、工作区、harness 和 scorer 做有无 skill 的配对试验,并以 ATIF 统一轨迹。145 个 skill 的扫描指标与 LLM judge 仅有 rho=0.14;947 个配对案例平均 composite Skill Lift 0.2134,72.8% 案例为正,说明静态文档门无法代替运行时测量。
–浏览

评论 · Comments