4 papers
From Atoms to Trees: Building a Structured Feature Forest with Hierarchical Sparse Autoencoders
Yifan Luo, Yang Zhan, Jiedong Jiang +4
Sparse autoencoders (SAEs) have proven effective for extracting monosemantic features from large language models (LLMs), yet these features are typically identified in isolation. H…
PDE Solvers Should Be Local: Fast, Stable Rollouts with Learned Local Stencils
Chun-Wun Cheng, Bin Dong, Carola-Bibiane Schönlieb +1
Neural operator models for solving partial differential equations (PDEs) often rely on global mixing mechanisms-such as spectral convolutions or attention-which tend to oversmooth…
CLGRPO: Reasoning Ability Enhancement for Small VLMs
Fanyi Wang, Binzhi Dong, Haotian Hu +2
Small Vision Language Models (SVLMs) generally refer to models with parameter sizes less than or equal to 2B. Their low cost and power consumption characteristics confer high comme…
Jailbreak Instruction-Tuned LLMs via end-of-sentence MLP Re-weighting
Yifan Luo, Zhennan Zhou, Meitan Wang +1
In this paper, we investigate the safety mechanisms of instruction fine-tuned large language models (LLMs). We discover that re-weighting MLP neurons can significantly compromise a…