6 papers
Col-Bandit: Query-Time Top- Estimation for Late-Interaction Retrieval
Roi Pony, Adi Raz Goldfarb, Oshri Naparstek +3
Multi-vector late-interaction retrievers such as ColBERT achieve state-of-the-art quality, but their query-time cost is dominated by exhaustively computing token-level MaxSim inter…
CARES: Context-Aware Resolution Selector for VLMs
Moshe Kimhi, Nimrod Shabtay, Raja Giryes +2
Large vision-language models (VLMs) commonly process images at native or high resolution to remain effective across tasks. This inflates visual tokens ofter to 97-99% of total toke…
Balanced Thinking: Improving Chain of Thought Training in Vision Language Models
Shaked Perek, Ben Wiesel, Avihu Dekel +2
Multimodal reasoning in vision-language models (VLMs) typically relies on a two-stage process: supervised fine-tuning (SFT) and reinforcement learning (RL). In standard SFT, all to…
Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs
Nimrod Shabtay, Moshe Kimhi, Artem Spector +5
Vision-language models (VLMs) typically process images at a native high-resolution, forcing a trade-off between accuracy and computational efficiency: high-resolution inputs captur…
WAVECLIP: Wavelet Tokenization for Adaptive-Resolution CLIP
Moshe Kimhi, Erez Koifman, Ehud Rivlin +2
We introduce WAVECLIP, a single unified model for adaptive resolution inference in CLIP, enabled by wavelet-based tokenization. WAVECLIP replaces standard patch embeddings with a m…
Advancing Speech Understanding in Speech-Aware Language Models with GRPO
Avishai Elmakies, Hagai Aronowitz, Nimrod Shabtay +3
In this paper, we introduce a Group Relative Policy Optimization (GRPO)-based method for training Speech-Aware Large Language Models (SALLMs) on open-format speech understanding ta…