论文把结构化输出中的 schema description 当作第二条指令通道,在两个厂商的十种配置上,用 nonce label 分类隔离标签先验。GPT-4.1 与无 reasoning 的 GPT-5.4 中,schema 放定义比 system prompt 低 11 - 13 点;当两者冲突时,错误 schema 可令准确率跌 5 - 45 点,Claude Haiku 4.5 更从 52.5% 掉到 7%。在 label 前增加必填 reasoning 字段却能把 schema-only 准确率抬高 15 - 24 点,说明 JSON Schema 不只是验证器配置,字段顺序和描述共同参与模型控制面。
–浏览
评论 · Comments