每日 Harness 开源 · Source
返回本期 · Back to 2026-07-15

论文 · Papers2026-07-15 · Wednesday, July 15, 2026

Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking

arxiv.org原文 ↗

评测方法基准其他垂直
Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking
这篇论文把 prompt wrapper 对 benchmark 的扰动量化为 FSI,并用 PSI 衡量输出是否还能被 schema 解析,而不是只报告模型分数。实验覆盖 14 个开源 LLM、25 种 wrapper 和 6 个 benchmark,显示 bias message、JSON schema、Markdown 等外层格式会改变 accuracy/F1 与 parseability。它值得放进评测流程讨论,因为它把“提示模板工程”从隐性变量变成了可审计的实验条件。
浏览

评论 · Comments