4 papers
HyperGS: Fast and Generalizable Gaussian Video Representation
Fatimah Zohra, Chen Zhao, Shuming Liu +2
Gaussian Splatting has emerged as an effective representation for video, but existing methods rely on per-video optimization. This leads to slow encoding and limits generalization…
Sparse Attention for Dense Open-Vocabulary Prediction in CLIP
Fatimah Zohra, Chen Zhao, Shuming Liu +1
Contrastive Language-Image Pre-training (CLIP) relies on softmax-based self-attention, a strictly positive distribution that assigns probability mass to every pair of tokens-even s…
-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment
Fatimah Zohra, Chen Zhao, Hani Itani +1
CLIP achieves strong zero-shot image-text retrieval by aligning global vision and text representations, yet it falls behind on fine-grained tasks even when fine-tuned on long, deta…
OpenTAD: A Unified Framework and Comprehensive Study of Temporal Action Detection
Shuming Liu, Chen Zhao, Fatimah Zohra +10
Temporal action detection (TAD) is a fundamental video understanding task that aims to identify human actions and localize their temporal boundaries in videos. Although this field…