Import AI 460: Reward hacking society, RSI data from Anthropic; and RL-based quadcopter racing
importai.substack.com原文 ↗
Import AI 本期串联三个研究信号:SocioHack 用 72 个 sandbox societal environments 测试制度性 reward hacking;Anthropic 内部数据显示 2026 年 merge 代码量相对 2021-2024 增长 8x;无人机多智能体 RL 在 22m/s+ 竞速中击败冠军级人类飞手,并把 collision rate 相对单智能体 baseline 降低 50%。
–浏览
评论 · Comments