collaborators

5 papers

cs.CV2026

CoCo-IR: Contextual Composed Image Retrieval

Shengcao Cao, Tanmaya Shekhar Dabral, Zhongli Ding +6

Current instruction-based image retrieval systems are powerful but limited to single-turn interactions, failing to capture the iterative nature of complex, real-world visual search…

cs.CV2026

Scaling Pre-training to One Hundred Billion Data for Vision Language Models

Xiao Wang, Ibrahim Alabdulmohsin, Daniel Salz +3

We provide an empirical investigation of the potential of pre-training vision-language models on an unprecedented scale: 100 billion examples. We find that model performance tends…

cs.CV2026

QAPruner: Quantization-Aware Vision Token Pruning for Multimodal Large Language Models

Xinhao Wang, Zhonyu Xia, Zhiwei Lin +2

Multimodal Large Language Models (MLLMs) have shown strong reasoning ability, but their high computational and memory costs hinder deployment in resource-constrained settings. Whil…

cs.CL2025

EmbeddingGemma: Powerful and Lightweight Text Representations

Henrique Schechter Vera, Sahil Dua, Biao Zhang +86

We introduce EmbeddingGemma, a new lightweight, open text embedding model based on the Gemma 3 language model family. Our innovative training recipe strategically captures knowledg…

cs.CL2025

Gemini Embedding: Generalizable Embeddings from Gemini

Jinhyuk Lee, Feiyang Chen, Sahil Dua +44

In this report, we introduce Gemini Embedding, a state-of-the-art embedding model leveraging the power of Gemini, Google's most capable large language model. Capitalizing on Gemini…