3 papers
cs.CV2025
WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
Shenghao Fu, Yukun Su, Fengyun Rao +3
Open-vocabulary object detection aims to detect arbitrary classes via text prompts. Methods without cross-modal fusion layers (non-fusion) offer faster inference by treating recogn…
cs.CV2025
WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens
Jian Yang, Dacheng Yin, Xiaoxuan He +6
Recent progress in multimodal large language models (MLLMs) has highlighted the challenge of efficiently bridging pre-trained Vision-Language Models (VLMs) with Diffusion Models. W…
cs.CV2025
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching
Xinli Yue, JianHui Sun, Junda Lu +6
With the rapid advancement of text-to-image (T2I) generation models, assessing the semantic alignment between generated images and text descriptions has become a significant resear…