4 citations · 6 across the 6 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
HarmoCLIP: Harmonizing Global and Regional Representations in Contrastive Vision-Language Models
Haoxi Zeng, Haoxuan Li, Yi Bin +4
Contrastive Language-Image Pre-training (CLIP) has demonstrated remarkable generalization ability and strong performance across a wide range of vision-language tasks. However, due…
cs.CV2025
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering
Jisheng Dang, Huilin Song, Junbin Xiao +6
Grounded Video Question Answering (Grounded VideoQA) requires aligning textual answers with explicit visual evidence. However, modern multimodal models often rely on linguistic pri…