15 papers
Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers
Chongjian Ge, Hanwen Jiang, Tianyu Wang +9
The paper presents Chimera, a hybrid visual diffusion transformer that processes text, image, and video tokens in a single raster-ordered stream using efficient attention mechanism…
Multi-Head Attention Residuals
Cheng Luo, Zefan Cai, Junjie Hu
Transformers propagate information across depth through a single additive residual stream: every sublayer reads only the most recent state. Attention residuals relax this by lettin…
Test-Time Training with Next-Token Prediction
Xuan Ouyang, Zefan Cai, Junjie Hu
Next-token prediction is the self-supervised signal that trains language models, and every observed prompt token provides the same signal at test time. We study whether this signal…
UniTemp: Unlocking Video Generation in Any Temporal Order via Bidirectional Distillation
Lin Zhang, Sicheng Mo, Zefan Cai +6
Autoregressive video diffusion models have emerged as a promising approach for long video generation, achieving strong performance in streaming settings. However, existing methods…
MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
Haozhe Zhao, Zefan Cai, Shuzheng Si +5
Recent text-to-image models produce high-quality results but still struggle with precise visual control, balancing multimodal inputs, and requiring extensive training for complex m…
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
Congmin Zheng, Jiachen Zhu, Jianghao Lin +6
Process Reward Models (PRMs) play a central role in evaluating and guiding multi-step reasoning in large language models (LLMs), especially for mathematical problem solving. Howeve…