2 papers
cs.CV2026
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No
Ji Huang, Barry Devereux, Hui Wang
Multimodal LLMs that recognise events reliably still fail to say when they happen. Prompted for timestamps, strong VLMs reach as little as [email protected] on Charades-STA, and t…
cs.CV2026
TBSG-Net: Temporal Bipartite Scene Graph Network for Fine-Grained Video Moment Retrieval
Ji Huang, Yongsheng Dai, Tianyu Ren +2
Recent advances in proposal-free Video Moment Retrieval (VMR) have highlighted the effectiveness of Static Scene Graphs (SSGs). By modeling objects and their relations at the frame…