PolyWorkBench 把长程 workplace agent 评测从单语设定推到多语言输入、推理、工具调用和结构化输出混在一起的场景。benchmark 包含 67 个任务、5 个领域:commerce、knowledge work、legal analysis、localization 和 manufacturing;评分混合 structural grading、executable verification 与 LLM semantic assessment。摘要中的经验结论是,state-of-the-art agent 在多语言 workflow 中相对单语 counterpart 明显退化,说明语言变化会和程序性决策步骤相互叠加。
–浏览
评论 · Comments