1 citations · 1 across the 2 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
Growing Visual Generative Capacity for Pre-Trained MLLMs
Hanyu Wang, Jiaming Han, Ziyan Yang +6
Multimodal large language models (MLLMs) extend the success of language models to visual understanding, and recent efforts have sought to build unified MLLMs that support both unde…
cs.CV2025
Long Context Tuning for Video Generation
Yuwei Guo, Ceyuan Yang, Ziyan Yang +5
Recent advances in video generation can produce realistic, minute-long single-shot videos with scalable diffusion transformers. However, real-world narrative videos require multi-s…
cs.CV2024
Parallelized Autoregressive Visual Generation
Yuqing Wang, Shuhuai Ren, Zhijie Lin +6
Autoregressive models have emerged as a powerful approach for visual generation but suffer from slow inference speed due to their sequential token-by-token prediction process. In t…