activity
20222026
most citedDiffusionBERT: Improving Generative Masked Language Models with Diffusion Models

13 citations · 27 across the 13 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2026

Tracing the Thought of a Grandmaster-level Chess-Playing Transformer

Rui Lin, Zhenyu Jin, Guancheng Zhou +7

While modern transformer neural networks achieve grandmaster-level performance in chess and other reasoning tasks, their internal computation process remains largely opaque. Focusi…

cs.LG2025

Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning

Junxuan Wang, Xuyang Ge, Wentao Shu +2

Transformer architectures, and their attention mechanisms in particular, form the foundation of modern large language models. While transformer models are widely believed to operat…

cs.LG2025

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition

Zhengfu He, Junxuan Wang, Rui Lin +5

We propose Low-Rank Sparse Attention (Lorsa), a sparse replacement model of Transformer attention layers to disentangle original Multi Head Self Attention (MHSA) into individually…

cs.LG2024★ 5 cited

Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders

Zhengfu He, Wentao Shu, Xuyang Ge +9

Sparse Autoencoders (SAEs) have emerged as a powerful unsupervised method for extracting sparse representations from language models, yet scalable training remains a significant ch…

cs.LG2024★ 2 cited

Automatically Identifying Local and Global Circuits with Linear Computation Graphs

Xuyang Ge, Fukang Zhu, Wentao Shu +3

Circuit analysis of any certain model behavior is a central task in mechanistic interpretability. We introduce our circuit discovery pipeline with Sparse Autoencoders (SAEs) and a…

cs.LG2024★ 1 cited

Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT

Zhengfu He, Xuyang Ge, Qiong Tang +3

Sparse dictionary learning has been a rapidly growing technique in mechanistic interpretability to attack superposition and extract more human-understandable features from model ac…