4 papers · 1 filter
IS-Diff: Improving Diffusion-Based Inpainting with Better Initial Seed
Yongzhe Lyu, Yu Wu, Yutian Lin +1
Diffusion models have shown promising results in free-form inpainting. Recent studies based on refined diffusion samplers or novel architectural designs led to realistic results an…
Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation
Feizhen Huang, Yu Wu, Yutian Lin +1
Video-to-Audio (V2A) Generation achieves significant progress and plays a crucial role in film and video post-production. However, current methods overlook the cinematic language,…
MAC: A Benchmark for Multiple Attributes Compositional Zero-Shot Learning
Shuo Xu, Sai Wang, Xinyue Hu +3
Compositional Zero-Shot Learning (CZSL) aims to learn semantic primitives (attributes and objects) from seen compositions and recognize unseen attribute-object compositions. Existi…
Improving Bird's Eye View Semantic Segmentation by Task Decomposition
Tianhao Zhao, Yongcan Chen, Yu Wu +8
Semantic segmentation in bird's eye view (BEV) plays a crucial role in autonomous driving. Previous methods usually follow an end-to-end pipeline, directly predicting the BEV segme…