每日 Harness 开源 · Source
返回本期 · Back to 2026-06-18

行业动态 · Industry News2026-06-18 · Thursday, June 18, 2026

Introducing LifeSciBench

openai.com原文 ↗

基准评测方法研究·科学
OpenAI 的 LifeSciBench 不是生物知识问答集,而是把应用生命科学研究拆成 evidence handling、analysis、design and optimization、scientific reasoning、validation and operations、translation、scientific communication 等 workflow。数据集规模是 750 个 expert-authored tasks、1,062 个 artifacts、173 位 scientist contributors、19,020 条 rubric criteria 和 453 位 expert reviewers;79% 任务需要多步推理或决策,53% 要处理 artifact。结果部分显示 GPT-Rosalind overall exact pass rate 从 GPT-5.5 的 25.7% 提到 36.1%,但 artifact-heavy 场景仍从 text-only 的 45.1% 降到 28.1%,说明真实科研可用性主要卡在证据、格式和操作约束。
浏览

评论 · Comments