6 papers
Demystifying Numerosity in Diffusion Models -- Limitations and Remedies
Yaqi Zhao, Xiaochen Wang, Li Dong +2
Numerosity remains a challenge for state-of-the-art text-to-image generation models like FLUX and GPT-4o, which often fail to accurately follow counting instructions in text prompt…
Kling-Foley: Multimodal Diffusion Transformer for High-Quality Video-to-Audio Generation
Jun Wang, Xijuan Zeng, Chunyu Qiang +19
We propose Kling-Foley, a large-scale multimodal Video-to-Audio generation model that synthesizes high-quality audio synchronized with video content. In Kling-Foley, we introduce m…
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
Yizhao Gao, Shuming Guo, Shijie Cao +12
We introduce SeerAttention-R, a sparse attention framework specifically tailored for the long decoding of reasoning models. Extended from SeerAttention, SeerAttention-R retains the…
Reinforcement Pre-Training
Qingxiu Dong, Li Dong, Yao Tang +4
In this work, we introduce Reinforcement Pre-Training (RPT) as a new scaling paradigm for large language models and reinforcement learning (RL). Specifically, we reframe next-token…
Rectified Sparse Attention
Yutao Sun, Tianzhu Ye, Li Dong +6
Efficient long-sequence generation is a critical challenge for Large Language Models. While recent sparse decoding methods improve efficiency, they suffer from KV cache misalignmen…
Probing the Visualization Literacy of Vision Language Models: the Good, the Bad, and the Ugly
Lianghan Dong, Anamaria Crisan
Vision Language Models (VLMs) demonstrate promising chart comprehension capabilities. Yet, prior explorations of their visualization literacy have been limited to assessing their r…