3 papers
cs.CV2025
InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search
Kaican Li, Lewei Yao, Jiannan Wu +7
The ability for AI agents to "think with images" requires a sophisticated blend of reasoning and perception. However, current open multimodal agents still largely fall short on the…
cs.CV2025
CannyEdit: Selective Canny Control and Dual-Prompt Guidance for Training-Free Image Editing
Weiyan Xie, Han Gao, Didan Deng +4
Recent advances in text-to-image (T2I) models have enabled training-free regional image editing by leveraging the generative priors of foundation models. However, existing methods…
cs.LG2024
Dual Risk Minimization: Towards Next-Level Robustness in Fine-tuning Zero-Shot Models
Kaican Li, Weiyan Xie, Yongxiang Huang +5
Fine-tuning foundation models often compromises their robustness to distribution shifts. To remedy this, most robust fine-tuning methods aim to preserve the pre-trained features. H…