4 papers
Where Culture Fades: Revealing the Cultural Gap in Text-to-Image Generation
Chuancheng Shi, Shangze Li, Shiming Guo +9
Multilingual text-to-image (T2I) models have advanced rapidly in terms of visual realism and semantic alignment, and are now widely utilized. Yet outputs vary across cultural conte…
DyST-XL: Dynamic Layout Planning and Content Control for Compositional Text-to-Video Generation
Weijie He, Mushui Liu, Yunlong Yu +2
Compositional text-to-video generation, which requires synthesizing dynamic scenes with multiple interacting entities and precise spatial-temporal relationships, remains a critical…
RestorerID: Towards Tuning-Free Face Restoration with ID Preservation
Jiacheng Ying, Mushui Liu, Zhe Wu +7
Blind face restoration has made great progress in producing high-quality and lifelike images. Yet it remains challenging to preserve the ID information especially when the degradat…
SAM 2: Segment Anything in Images and Videos
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu +15
We present Segment Anything Model 2 (SAM 2), a foundation model towards solving promptable visual segmentation in images and videos. We build a data engine, which improves model an…