每日 Harness 开源 · Source
返回本期 · Back to 2026-06-18

论文 · Papers2026-06-18 · Thursday, June 18, 2026

How Inference Compute Shapes Frontier LLM Evaluation

arxiv.org原文 ↗

How Inference Compute Shapes Frontier LLM Evaluation
这篇论文把 benchmark score 视为模型能力与测试时计算协议的共同产物,而不是模型的单点属性。作者在最多 12 个 frontier language model 和 7 个覆盖软件工程、数学、医学、网络安全的 benchmark 上,组合更大 token budget、context compaction、repeated submission attempts,以及模型自导或 minimal correctness feedback。结果显示大预算能显著改善多个领域表现,固定预算会越来越低估新一代模型;安全或政策相关评测若只报一个 restrictive budget,结论会混入协议选择偏差。
浏览

评论 · Comments