5 papers
Scaling Parallel Sequence Models to Foundation-Scale Vision Encoders
Yitong Jiang, Hongjun Wang, Collin McCarthy +15
Vision foundation models are bottlenecked by the quadratic cost of self-attention, which limits usable resolution and increases the cost of large-scale pretraining. Subquadratic al…
WonderVerse: Extendable 3D Scene Generation with Video Generative Models
Hao Feng, Zhi Zuo, Jia-Hui Pan +4
We introduce \textit{WonderVerse}, a simple but effective framework for generating extendable 3D scenes. Unlike existing methods that rely on iterative depth estimation and image i…
CubeComposer: Spatio-Temporal Autoregressive 4K 360° Video Generation from Perspective Video
Lingen Li, Guangzhi Wang, Xiaoyu Li +5
Generating high-quality 360° panoramic videos from perspective input is one of the crucial applications for virtual reality (VR), whereby high-resolution videos are especially imp…
CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects
Zhao Wang, Aoxue Li, Lingting Zhu +3
Customized text-to-video generation aims to generate high-quality videos guided by text prompts and subject references. Current approaches for personalizing text-to-video generatio…
AnyCharV: Bootstrap Controllable Character Video Generation with Fine-to-Coarse Guidance
Zhao Wang, Hao Wen, Lingting Zhu +3
Character video generation is a significant real-world application focused on producing high-quality videos featuring specific characters. Recent advancements have introduced vario…