10 papers
Krause Synchronization Transformers
Jingkun Liu, Yisong Yue, Max Welling +1
Self-attention in Transformers relies on globally normalized softmax weights, causing all tokens to compete for influence at every layer. When composed across depth, this interacti…
Spontaneous symmetry breaking and Goldstone modes for deep information propagation
Nabil Iqbal, T. Anderson Keller, Yue Song +2
In physical systems, whenever a continuous symmetry is spontaneously broken, the system possesses excitations called Goldstone modes, which allow coherent information propagation o…
Higher-Order Equilibrium Tracking for EM-Compressible Online Estimation
ZhiMing Li, Yue Song
We study online estimation in latent-variable models by recasting the problem as tracking a moving empirical equilibrium. Standard online EM and stochastic approximation analyses p…
Anisotropic Modality Align
Xiaomin Yu, Yijiang Li, Yuhui Zhang +8
Training multimodal large language models has long been limited by the scarcity of high-quality paired multimodal data. Recent studies show that the shared representation space of…
MOVA: Towards Scalable and Synchronized Video-Audio Generation
OpenMOSS Team, Donghua Yu, Mingshu Chen +38
Audio is indispensable for real-world video, yet generation models have largely overlooked audio components. Current approaches to producing audio-visual content often rely on casc…
Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks
Hoagy Cunningham, Jerry Wei, Zihan Wang +26
We introduce enhanced Constitutional Classifiers that deliver production-grade jailbreak robustness with dramatically reduced computational costs and refusal rates compared to prev…