From the 2 of 11 linked papers with an AI index.
11 papers
HyperGS: Fast and Generalizable Gaussian Video Representation
Fatimah Zohra, Chen Zhao, Shuming Liu +2
HyperGS is a feedforward model that predicts Gaussian video representations directly from input video in a single forward pass, achieving orders-of-magnitude faster encoding and be…
Sparse Attention for Dense Open-Vocabulary Prediction in CLIP
Fatimah Zohra, Chen Zhao, Shuming Liu +1
The paper replaces the softmax in CLIP's visual self‑attention with the α‑entmax transform to create sparse attention, which reduces noise from irrelevant tokens and improves dense…
GoTTA be Diverse: Rethinking Memory Policies for Test-Time Adaptation
Shyma Alhuwaider, Yasmeen Alsaedy, Merey Ramazanova +2
Test-time adaptation (TTA) enables a pre-trained model to adapt online to an unlabeled test stream under distribution shift. While most TTA research focuses on the adaptation objec…
CounterCount: A Diagnostic Framework for Counting Bias in Vision Language Models
Reem Alzahrani, Hassan Alshanqiti, Bushra Bin Hemid +3
Vision-Language Models (VLMs) excel at multimodal reasoning, yet it remains unclear whether their answers are grounded in visual evidence or driven by learned language and world pr…
-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment
Fatimah Zohra, Chen Zhao, Hani Itani +1
CLIP achieves strong zero-shot image-text retrieval by aligning global vision and text representations, yet it falls behind on fine-grained tasks even when fine-tuned on long, deta…
SEVERE++: Evaluating Benchmark Sensitivity in Generalization of Video Representation Learning
Fida Mohammad Thoker, Letian Jiang, Chen Zhao +4
Continued advances in self-supervised learning have led to significant progress in video representation learning, offering a scalable alternative to supervised approaches by removi…