Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment
Peiyuan Zhu, Shaoan Xie, Zijian Li +5
Contrastive pre-training has propelled video-text alignment, yet models often inherit the critical limitations of their image-text predecessors like CLIP, resulting in entangled re…
cs.CV2026
Multimodal LLMs under Pairwise Modalities
Yan Li, Yunlong Deng, Yuewen Sun +3
Despite the impressive results achieved by multimodal large language models (MLLMs), their training typically relies on jointly curated multimodal data, requiring substantial human…
cs.CV2024
Learning Causal Domain-Invariant Temporal Dynamics for Few-Shot Action Recognition
Yuke Li, Guangyi Chen, Ben Abramowitz +2
Few-shot action recognition aims at quickly adapting a pre-trained model to the novel data with a distribution shift using only a limited number of samples. Key challenges include…