2 papers
cs.CV2026
HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
Joungbin An, Kristen Grauman
Video temporal grounding, the task of localizing the start and end times of a natural language query in untrimmed video, requires capturing both global context and fine-grained tem…
cs.CV2025
Progress-Aware Video Frame Captioning
Zihui Xue, Joungbin An, Xitong Yang +1
While image captioning provides isolated descriptions for individual images, and video captioning offers one single narrative for an entire video clip, our work explores an importa…