2 papers
cs.CV2026
Continual Video-MLLM Adaptation over Evolving Domains
Rui Cheng, Meixing Shi, Yuxiang Cai +3
Video multimodal large language models have shown strong capability in video understanding, yet their adaptation to sequentially evolving domains remains underexplored. In real-wor…
cs.CV2025
MVAM: Multi-View Attention Method for Fine-grained Image-Text Matching
Wanqing Cui, Rui Cheng, Jiafeng Guo +1
Existing two-stream models, such as CLIP, encode images and text through independent representations, showing good performance while ensuring retrieval speed, have attracted attent…