6 papers
Emergent Hierarchical Monosemantic Neurons from the Group-Contrastive Forward-Forward Algorithm
Yiming Tang, Qinglin Qi, Zhaoqian Yao +2
Mechanistic interpretability has made significant strides in understanding neural network representations, with sparse dictionary learning (SDL) methods, most prominently sparse au…
SAHG: Sector-Anisotropic Hyperbolic Graph Model for Social Bot Detection
Hanning Lu, Yingguang Yang, Jinwei Su +8
LLM-driven social bots can generate fluent, human-like text, reducing the discriminative advantage of content-based detection alone. However, coordinated campaigns still leave rela…
A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima
Yiming Tang, Harshvardhan Saini, Zhaoqian Yao +6
As AI models achieve remarkable capabilities across diverse domains, understanding what representations they learn and how they encode concepts has become increasingly important fo…
MagnifierSketch: Quantile Estimation Centered at One Point
Jiarui Guo, Qiushi Lyu, Yuhan Wu +6
In this paper, we take into consideration quantile estimation in data stream models, where every item in the data stream is a key-value pair. Researchers sometimes aim to estimate…
How Syntax Specialization Emerges in Language Models
Xufeng Duan, Zhaoqian Yao, Yunhao Zhang +2
Large language models (LLMs) have been found to develop surprising internal specializations: Individual neurons, attention heads, and circuits become selectively sensitive to synta…
QET: Enhancing Quantized LLM Parameters and KV cache Compression through Element Substitution and Residual Clustering
Yanshu Wang, Wang Li, Zhaoqian Yao +1
The matrix quantization entails representing matrix elements in a more space-efficient form to reduce storage usage, with dequantization restoring the original matrix for use. We f…