5 papers · 1 filter
Comp-Attn: Present-and-Align Attention for Compositional Video Generation
Hongyu Zhang, Yufan Deng, Shenghai Yuan +5
In the domain of text-to-video (T2V) generation, reliably synthesizing compositional content involving multiple subjects with intricate relations is still underexplored. The main c…
Local Action-Guided Motion Diffusion Model for Text-to-Motion Generation
Peng Jin, Hao Li, Zesen Cheng +6
Text-to-motion generation requires not only grounding local actions in language but also seamlessly blending these individual actions to synthesize diverse and realistic global mot…
GraCo: Granularity-Controllable Interactive Segmentation
Yian Zhao, Kehan Li, Zesen Cheng +6
Interactive Segmentation (IS) segments specific objects or parts in the image according to user input. Current IS pipelines fall into two categories: single-granularity output and…
Instance Brownian Bridge as Texts for Open-vocabulary Video Instance Segmentation
Zesen Cheng, Kehan Li, Hao Li +5
Temporally locating objects with arbitrary class texts is the primary pursuit of open-vocabulary Video Instance Segmentation (VIS). Because of the insufficient vocabulary of video…
Towards Real-World Burst Image Super-Resolution: Benchmark and Method
Pengxu Wei, Yujing Sun, Xingbei Guo +4
Despite substantial advances, single-image super-resolution (SISR) is always in a dilemma to reconstruct high-quality images with limited information from one input image, especial…