collaborators

6 papers

cs.CL2025

Steering Information Utility in Key-Value Memory for Language Model Post-Training

Chunyuan Deng, Ruidi Chang, Hanjie Chen

Recent advancements in language models (LMs) have marked a shift toward the growing importance of post-training. Yet, post-training approaches such as supervised fine-tuning (SFT)…

cs.CL2025

Learning Distribution-Wise Control in Representation Space for Language Models

Chunyuan Deng, Ruidi Chang, Hanjie Chen

Interventions in language models (LMs) are applied strategically to steer model behavior during the forward pass. Learnable interventions, also known as representation fine-tuning,…

cs.LG2025

SAFR: Neuron Redistribution for Interpretability

Ruidi Chang, Chunyuan Deng, Hanjie Chen

Superposition refers to encoding representations of multiple features within a single neuron, which is common in deep neural networks. This property allows neurons to combine and r…

cs.AI2025

Rethinking Diverse Human Preference Learning through Principal Component Analysis

Feng Luo, Rui Yang, Hao Sun +5

Understanding human preferences is crucial for improving foundation models and building personalized AI systems. However, preferences are inherently diverse and complex, making it…

cs.CL2025

Knowing When to Stop: Efficient Context Processing via Latent Sufficiency Signals

Roy Xie, Junlin Wang, Paul Rosu +4

Large language models (LLMs) process entire input contexts indiscriminately, which is inefficient when the information required to answer a query is localized within the context. W…

cs.LG2024

Language Models are Symbolic Learners in Arithmetic

Chunyuan Deng, Zhiqi Li, Roy Xie +2

The prevailing question in LM performing arithmetic is whether these models learn to truly compute or if they simply master superficial pattern matching. In this paper, we argues f…