26 citations · 87 across the 35 of their papers we have counts for
5 papers · 1 filter
TimeLogic Challenge @ CVPR 2026: Strong MLLMs Meet Evidence-Seeking Agents for Temporal-Logic Video Question Answering
Zhaoyang Xu, Xusheng He, Wei Liu +2
Temporal-logic video question answering requires a model to reason about when actions occur relative to one another, such as before, after, until, since, overlap, and multi-event c…
FineBadminton: A Multi-Level Dataset for Fine-Grained Badminton Video Understanding
Xusheng He, Wei Liu, Shanshan Ma +3
Fine-grained analysis of complex and high-speed sports like badminton presents a significant challenge for Multimodal Large Language Models (MLLMs), despite their notable advanceme…
RA-BLIP: Multimodal Adaptive Retrieval-Augmented Bootstrapping Language-Image Pre-training
Muhe Ding, Yang Ma, Pengda Qin +3
Multimodal Large Language Models (MLLMs) have recently received substantial interest, which shows their emerging potential as general-purpose models for various vision-language tas…
Self-Training Boosted Multi-Factor Matching Network for Composed Image Retrieval
Haokun Wen, Xuemeng Song, Jianhua Yin +3
The composed image retrieval (CIR) task aims to retrieve the desired target image for a given multimodal query, i.e., a reference image with its corresponding modification text. Th…
Micro-video Tagging via Jointly Modeling Social Influence and Tag Relation
Xiao Wang, Tian Gan, Yinwei Wei +3
The last decade has witnessed the proliferation of micro-videos on various user-generated content platforms. According to our statistics, around 85.7\% of micro-videos lack annotat…