6 citations · 6 across the 3 of their papers we have counts for
4 papers
T2VParser: Adaptive Decomposition Tokens for Partial Alignment in Text to Video Retrieval
Yili Li, Gang Xiong, Gaopeng Gou +4
Text-to-video retrieval essentially aims to train models to align visual content with textual descriptions accurately. Due to the impressive general multimodal knowledge demonstrat…
ProAPO: Progressively Automatic Prompt Optimization for Visual Classification
Xiangyan Qu, Gaopeng Gou, Jiamin Zhuang +5
Vision-language models (VLMs) have made significant progress in image classification by training with large-scale paired image-text data. Their performances largely depend on the p…
T2VIndexer: A Generative Video Indexer for Efficient Text-Video Retrieval
Yili Li, Jing Yu, Keke Gai +3
Current text-video retrieval methods mainly rely on cross-modal matching between queries and videos to calculate their similarity scores, which are then sorted to obtain retrieval…
IIU: Independent Inference Units for Knowledge-based Visual Question Answering
Yili Li, Jing Yu, Keke Gai +1
Knowledge-based visual question answering requires external knowledge beyond visible content to answer the question correctly. One limitation of existing methods is that they focus…