4 papers · 1 filter
ProAPO: Progressively Automatic Prompt Optimization for Visual Classification
Xiangyan Qu, Gaopeng Gou, Jiamin Zhuang +5
Vision-language models (VLMs) have made significant progress in image classification by training with large-scale paired image-text data. Their performances largely depend on the p…
T2VParser: Adaptive Decomposition Tokens for Partial Alignment in Text to Video Retrieval
Yili Li, Gang Xiong, Gaopeng Gou +4
Text-to-video retrieval essentially aims to train models to align visual content with textual descriptions accurately. Due to the impressive general multimodal knowledge demonstrat…
MADS: Multi-Attribute Document Supervision for Zero-Shot Image Classification
Xiangyan Qu, Jing Yu, Jiamin Zhuang +3
Zero-shot learning (ZSL) aims to train a model on seen classes and recognize unseen classes by knowledge transfer through shared auxiliary information. Recent studies reveal that d…
Visual-Semantic Decomposition and Partial Alignment for Document-based Zero-Shot Learning
Xiangyan Qu, Jing Yu, Keke Gai +5
Recent work shows that documents from encyclopedias serve as helpful auxiliary information for zero-shot learning. Existing methods align the entire semantics of a document with co…