3 papers
cs.CV2025
Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding
Zilin Du, Haoxin Li, Jianfei Yu +1
Visual grounding aims to localize the image regions based on a textual query. Given the difficulty of large-scale data curation, we investigate how to effectively learn visual grou…
cs.CV2025
Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic Data
Haoxin Li, Boyang Li
Paired image-text data with subtle variations in-between (e.g., people holding surfboards vs. people holding shovels) hold the promise of producing Vision-Language Models with prop…
cs.CV2025
Learning to Animate Images from A Few Videos to Portray Delicate Human Actions
Haoxin Li, Yingchen Yu, Qilong Wu +3
Despite recent progress, video generative models still struggle to animate static images into videos that portray delicate human actions, particularly when handling uncommon or nov…