每日 Harness 开源 · Source
返回本期 · Back to 2026-06-19

论文 · Papers2026-06-19 · Friday, June 19, 2026

Bag of Dims: Training-Free Mechanistic Interpretability via Dimension-Level Sign Patterns

arxiv.org原文 ↗

Bag of Dims: Training-Free Mechanistic Interpretability via Dimension-Level Sign Patterns
论文认为 transformer hidden state 的标准维度本身就可作为训练自由的特征基,符号编码语义,幅值编码置信度。作者在 Qwen 3.5-4B、Gemma 3-4B、Mistral 7B 上发现只保留 sign pattern 仍能达到 72%-93% top-5 next-token accuracy,无监督还可发现 1500 个特征并保持 99% sparsity。它对 mechanistic interpretability 的冲击在于:不一定先训练 SAE 才能读出结构。
浏览

评论 · Comments