Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
arxiv.org原文 ↗
Nemotron 3 Ultra 是 550B total、55B active 的 MoE hybrid Mamba-Attention 模型。NVIDIA 报告其用 20T text tokens 预训练,随后扩展到 1M context,并经过 SFT、RL 与 MOPD post-training。LatentMoE、MTP、NVFP4 pre-training、multi-environment RLVR 和 reasoning budget control 共同服务于长程 agentic workload;官方称相对公开 SOTA LLM 在同等准确率下最高有约 6x 推理吞吐。
–浏览
评论 · Comments