Kimi K3, and what we can still learn from the pelican benchmark
simonwillison.net原文 ↗
Simon Willison 的文章实际围绕 Kimi K2 与 pelican benchmark 展开,文中提到 Kimi K2 是 1T 参数、约 32B active parameters 的 MoE 模型。作者没有把基准当成最终排名,而是用一个具体表现案例分析模型在图像指令任务中的行为。它提醒读者:小基准仍有价值,但价值来自诊断模型如何思考,而不是把分数截图当结论。
–浏览
评论 · Comments