Bag of Dims: Training-Free Mechanistic Interpretability via Dimension-Level Sign Patterns
arxiv.org原文 ↗
论文认为 transformer hidden state 的标准维度本身就可作为训练自由的特征基,符号编码语义,幅值编码置信度。作者在 Qwen 3.5-4B、Gemma 3-4B、Mistral 7B 上发现只保留 sign pattern 仍能达到 72%-93% top-5 next-token accuracy,无监督还可发现 1500 个特征并保持 99% sparsity。它对 mechanistic interpretability 的冲击在于:不一定先训练 SAE 才能读出结构。
–浏览
评论 · Comments