3 papers
cs.CV2026
UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
Hao Tang, Chenwei Xie, Xiaoyi Bao +4
In this paper, we propose UniLIP, a unified framework that adapts CLIP for multimodal understanding, generation and editing. Although CLIP excels at understanding, it lacks reconst…
cs.CV2025
UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface
Hao Tang, Chenwei Xie, Haiyang Wang +5
Generalist models have achieved remarkable success in both language and vision-language tasks, showcasing the potential of unified modeling. However, effectively integrating fine-g…
cs.CV2025
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding
Xiaoyi Bao, Chenwei Xie, Hao Tang +4
In recent years, the introduction of Multi-modal Large Language Models (MLLMs) into video understanding tasks has become increasingly prevalent. However, how to effectively integra…