collaborators

15 papers

cs.CV2026

Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers

Chongjian Ge, Hanwen Jiang, Tianyu Wang +9

The paper presents Chimera, a hybrid visual diffusion transformer that processes text, image, and video tokens in a single raster-ordered stream using efficient attention mechanism…

cs.AI2026

Multi-Head Attention Residuals

Cheng Luo, Zefan Cai, Junjie Hu

Transformers propagate information across depth through a single additive residual stream: every sublayer reads only the most recent state. Attention residuals relax this by lettin…

cs.CL2026

Test-Time Training with Next-Token Prediction

Xuan Ouyang, Zefan Cai, Junjie Hu

Next-token prediction is the self-supervised signal that trains language models, and every observed prompt token provides the same signal at test time. We study whether this signal…

cs.CV2026

UniTemp: Unlocking Video Generation in Any Temporal Order via Bidirectional Distillation

Lin Zhang, Sicheng Mo, Zefan Cai +6

Autoregressive video diffusion models have emerged as a promising approach for long video generation, achieving strong performance in streaming settings. However, existing methods…

cs.CV2026

MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models

Haozhe Zhao, Zefan Cai, Shuzheng Si +5

Recent text-to-image models produce high-quality results but still struggle with precise visual control, balancing multimodal inputs, and requiring extensive training for complex m…

cs.CL2026

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning

Congmin Zheng, Jiachen Zhu, Jianghao Lin +6

Process Reward Models (PRMs) play a central role in evaluating and guiding multi-step reasoning in large language models (LLMs), especially for mathematical problem solving. Howeve…