From the 2 of 5 linked papers with an AI index.
5 papers
HyperGS: Fast and Generalizable Gaussian Video Representation
Fatimah Zohra, Chen Zhao, Shuming Liu +2
HyperGS is a feedforward model that predicts Gaussian video representations directly from input video in a single forward pass, achieving orders-of-magnitude faster encoding and be…
Sparse Attention for Dense Open-Vocabulary Prediction in CLIP
Fatimah Zohra, Chen Zhao, Shuming Liu +1
The paper replaces the softmax in CLIP's visual self‑attention with the α‑entmax transform to create sparse attention, which reduces noise from irrelevant tokens and improves dense…
GenHOI: Towards Object-Consistent Hand-Object Interaction with Temporally Balanced and Spatially Selective Object Injection
Xuan Huang, Mochu Xiang, Zhelun Shen +9
Hand-Object Interaction (HOI) remains a core challenge in digital human video synthesis, where models must generate physically plausible contact and preserve object identity across…
-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment
Fatimah Zohra, Chen Zhao, Hani Itani +1
CLIP achieves strong zero-shot image-text retrieval by aligning global vision and text representations, yet it falls behind on fine-grained tasks even when fine-tuned on long, deta…
OpenTAD: A Unified Framework and Comprehensive Study of Temporal Action Detection
Shuming Liu, Chen Zhao, Fatimah Zohra +10
Temporal action detection (TAD) is a fundamental video understanding task that aims to identify human actions and localize their temporal boundaries in videos. Although this field…