3 papers
cs.CV2025
Rethinking The Training And Evaluation of Rich-Context Layout-to-Image Generation
Jiaxin Cheng, Zixu Zhao, Tong He +3
Recent advancements in generative models have significantly enhanced their capacity for image generation, enabling a wide range of applications such as image editing, completion an…
cs.CV2024
VideoSAM: Open-World Video Segmentation
Pinxue Guo, Zixu Zhao, Jianxiong Gao +5
Video segmentation is essential for advancing robotics and autonomous driving, particularly in open-world settings where continuous perception and object association across video f…
cs.CV2024
Unsupervised Open-Vocabulary Object Localization in Videos
Ke Fan, Zechen Bai, Tianjun Xiao +11
In this paper, we show that recent advances in video representation learning and pre-trained vision-language models allow for substantial improvements in self-supervised video obje…