每日 Harness 开源 · Source
返回本期 · Back to 2026-09-11

博客文章 · Blog Posts2026-09-11 · Friday, September 11, 2026

GPT-6 Astra, looped transformers, and hidden reasoning

magazine.sebastianraschka.com原文 ↗

GPT-6 Astra, looped transformers, and hidden reasoning
Raschka 解释 looped transformer 通过重复同一组 block 增加有效深度,并串联 Universal Transformer、Ouro 与 Mixture-of-Recursions 的路由思路。文中以 Nanbeige 为例:22 个 block 循环两次得到 44 次应用,参数存储约减半,但前向计算和 KV cache 仍接近 44 层模型,说明省内存不等于省算力。
–浏览

评论 · Comments