collaborators

10 papers

cs.CV2026

VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization

Junhao Cheng, Liang Hou, Tianxiong Zhong +4

The recent "Reasoning with Video" paradigm utilizes Video Generation Models (VGMs) to generate temporally coherent visual trajectories to complete reasoning tasks. Although state-o…

cs.CV2026

Stable Velocity: A Variance Perspective on Flow Matching

Donglin Yang, Yongxing Zhang, Xin Yu +5

While flow matching is elegant, its reliance on single-sample conditional velocities leads to high-variance training targets that destabilize optimization and slow convergence. By…

cs.GR2026

SURF: Signature-Retained Fast Video Generation

Kaixin Ding, Xi Chen, Sihui Ji +4

The demand for high-resolution video generation is growing rapidly. However, the generation resolution is severely constrained by slow inference speeds. For instance, Wan2.1 requir…

cs.CV2026

RectifiedHR: Enable Efficient High-Resolution Synthesis via Energy Rectification

Zhen Yang, Guibao Shen, Minyang Li +5

Diffusion models have achieved remarkable progress across various visual generation tasks. However, their performance significantly declines when generating content at resolutions…

cs.CV2026

VMonarch: Efficient Video Diffusion Transformers with Structured Attention

Cheng Liang, Haoxian Chen, Liang Hou +4

The quadratic complexity of the attention mechanism severely limits the context scalability of Video Diffusion Transformers (DiTs). We find that the highly sparse spatio-temporal a…

cs.CV2026

Boosting Resolution Generalization of Diffusion Transformers with Randomized Positional Encodings

Liang Hou, Cong Liu, Mingwu Zheng +4

Resolution generalization in image generation tasks enables the production of higher-resolution images with lower training resolution overhead. However, a key obstacle for diffusio…