Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
AdsQA: Towards Advertisement Video Understanding
Xinwei Long, Kai Tian, Peng Xu +10
Large language models (LLMs) have taken a great step towards AGI. Meanwhile, an increasing number of domain-specific problems such as math and programming boost these general-purpo…
cs.CV2025
Describe Anything in Medical Images
Xi Xiao, Yunbei Zhang, Thanh-Huy Nguyen +10
Localized image captioning has made significant progress with models like the Describe Anything Model (DAM), which can generate detailed region-specific descriptions without explic…
cs.CV2025
CIBR: Cross-modal Information Bottleneck Regularization for Robust CLIP Generalization
Yingrui Ji, Xi Xiao, Gaofei Chen +5
Contrastive Language-Image Pretraining (CLIP) has achieved remarkable success in cross-modal tasks such as zero-shot image classification and text-image retrieval by effectively al…