works on

From the 2 of 11 linked papers with an AI index.

collaborators

11 papers

cs.CV2026

HyperGS: Fast and Generalizable Gaussian Video Representation

Fatimah Zohra, Chen Zhao, Shuming Liu +2

HyperGS is a feedforward model that predicts Gaussian video representations directly from input video in a single forward pass, achieving orders-of-magnitude faster encoding and be…

cs.CV2026

Sparse Attention for Dense Open-Vocabulary Prediction in CLIP

Fatimah Zohra, Chen Zhao, Shuming Liu +1

The paper replaces the softmax in CLIP's visual self‑attention with the α‑entmax transform to create sparse attention, which reduces noise from irrelevant tokens and improves dense…

cs.CV2026

GoTTA be Diverse: Rethinking Memory Policies for Test-Time Adaptation

Shyma Alhuwaider, Yasmeen Alsaedy, Merey Ramazanova +2

Test-time adaptation (TTA) enables a pre-trained model to adapt online to an unlabeled test stream under distribution shift. While most TTA research focuses on the adaptation objec…

cs.CV2026

CounterCount: A Diagnostic Framework for Counting Bias in Vision Language Models

Reem Alzahrani, Hassan Alshanqiti, Bushra Bin Hemid +3

Vision-Language Models (VLMs) excel at multimodal reasoning, yet it remains unclear whether their answers are grounded in visual evidence or driven by learned language and world pr…

cs.CV2026

-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment

Fatimah Zohra, Chen Zhao, Hani Itani +1

CLIP achieves strong zero-shot image-text retrieval by aligning global vision and text representations, yet it falls behind on fine-grained tasks even when fine-tuned on long, deta…

cs.CV2025

SEVERE++: Evaluating Benchmark Sensitivity in Generalization of Video Representation Learning

Fida Mohammad Thoker, Letian Jiang, Chen Zhao +4

Continued advances in self-supervised learning have led to significant progress in video representation learning, offering a scalable alternative to supervised approaches by removi…