13 papers
The Illusion of Specialization: Unveiling the Domain-Invariant "Standing Committee" in Mixture-of-Experts Models
Yan Wang, Yitao Xu, Nanhan Shen +3
Mixture of Experts models are widely assumed to achieve domain specialization through sparse routing. In this work, we question this assumption by introducing COMMITTEEAUDIT, a pos…
Exploring Concept Subspace for Self-explainable Text-Attributed Graph Learning
Xiaoxue Han, Libo Zhang, Zining Zhu +1
We introduce Graph Concept Bottleneck (GCB) as a new paradigm for self-explainable text-attributed graph learning. GCB maps graphs into a subspace, concept bottleneck, where each c…
Feature-Guided SAE Steering for Refusal-Rate Control using Contrasting Prompts
Samaksh Bhargav, Zining Zhu
Large Language Model (LLM) deployment requires guiding the LLM to recognize and not answer unsafe prompts while complying with safe prompts. Previous methods for achieving this req…
Sheaf Discovery with Joint Computation Graph Pruning and Flexible Granularity
Lei Yu, Jingcheng Niu, Zining Zhu +2
In this paper, we introduce DiscoGP, a novel framework for extracting self-contained modular units, or sheaves, within neural language models (LMs). Sheaves extend the concept of f…
Distribution Prompting: Understanding the Expressivity of Language Models Through the Next-Token Distributions They Can Produce
Haojin Wang, Zining Zhu, Freda Shi
Autoregressive neural language models (LMs) generate a probability distribution over tokens at each time step given a prompt. In this work, we attempt to systematically understand…
How Well Can Knowledge Edit Methods Edit Perplexing Knowledge?
Huaizhi Ge, Frank Rudzicz, Zining Zhu
Large language models (LLMs) have demonstrated remarkable capabilities, but updating their knowledge post-training remains a critical challenge. While recent model editing techniqu…