2 citations · 4 across the 6 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No
Ji Huang, Barry Devereux, Hui Wang
Multimodal LLMs that recognise events reliably still fail to say when they happen. Prompted for timestamps, strong VLMs reach as little as [email protected] on Charades-STA, and t…
cs.CV2024
Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
Jiajun Fei, Dian Li, Zhidong Deng +3
Multi-modal large language models (MLLMs) have demonstrated considerable potential across various downstream tasks that require cross-domain knowledge. MLLMs capable of processing…