activity
20242026
collaborators
Showing cs.CVShow all

11 papers · 1 filter

cs.CV2026

Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model

Xinyin Ma, Julius Berner, Chao Liu +3

Recent progress in large-scale generative models has substantially advanced video generation, yet existing methods remain constrained by a rigid inference paradigm. Bidirectional d…

cs.CV2026

Q-ARVD: Quantizing Autoregressive Video Diffusion Models

Siao Tang, Xinyin Ma, Gongfan Fang +2

Autoregressive video diffusion models (ARVDs) have emerged as a promising architecture for streaming video generation, paving the way for real-time interactive video generation and…

cs.CV2025

In-Video Instructions: Visual Signals as Generative Control

Gongfan Fang, Xinyin Ma, Xinchao Wang

Large-scale video generative models have recently demonstrated strong visual capabilities, enabling the prediction of future frames that adhere to the logical and physical cues in…

cs.CV2025

Dimple: Discrete Diffusion Multimodal Large Language Model with Parallel Decoding

Runpeng Yu, Xinyin Ma, Xinchao Wang

In this work, we propose Dimple, the first Discrete Diffusion Multimodal Large Language Model (DMLLM). We observe that training with a purely discrete diffusion approach leads to s…

cs.CV2025

Introducing Visual Perception Token into Multimodal Large Language Model

Runpeng Yu, Xinyin Ma, Xinchao Wang

To utilize visual information, Multimodal Large Language Model (MLLM) relies on the perception process of its vision encoder. The completeness and accuracy of visual perception sig…

cs.CV2024

Remix-DiT: Mixing Diffusion Transformers for Multi-Expert Denoising

Gongfan Fang, Xinyin Ma, Xinchao Wang

Transformer-based diffusion models have achieved significant advancements across a variety of generative tasks. However, producing high-quality outputs typically necessitates large…