2 papers
cs.CV2025
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding
Yun Li, Zhe Liu, Yajing Kong +6
Applying Multimodal Large Language Models (MLLMs) to video understanding presents significant challenges due to the need to model temporal relations across frames. Existing approac…
cs.CV2024
Compositional Zero-Shot Learning with Contextualized Cues and Adaptive Contrastive Training
Yun Li, Zhe Liu, Lina Yao
Compositional Zero-Shot Learning (CZSL) aims to recognize unseen combinations of seen attributes and objects. Current CLIP-based methods in CZSL, despite their advancements, often…