4 papers
Adaptive Multilevel Twisted Sequential Monte Carlo for Rare Events Estimation in Language Models
Zixuan Liu, Fangzheng Wu, Brian Summa +1
Rare unsafe behaviors in large language models can remain practically significant even when their probability is extremely small, particularly at deployment scales involving millio…
Robust General Utility for Reinforcement Learning
Zixuan Liu, Fangzheng Wu, Brian Summa +1
Reinforcement learning (RL) with general utility extends classic RL by optimizing an arbitrary utility functional of the policy-induced occupancy measure, thereby enabling a broade…
Attention Sinks in Diffusion Transformers: A Causal Analysis
Fangzheng Wu, Brian Summa
Attention sinks -- tokens that receive disproportionate attention mass -- are assumed to be functionally important in autoregressive language models, but their role in diffusion tr…
Model-Centric Diagnostics: A Framework for Internal State Readouts
Fangzheng Wu, Brian Summa
We present a model-centric diagnostic framework that treats training state as a latent variable and unifies a family of internal readouts -- head-gradient norms, confidence, entrop…