1 citations · 1 across the 4 of their papers we have counts for
7 papers
Balanced Thinking: Improving Chain of Thought Training in Vision Language Models
Shaked Perek, Ben Wiesel, Avihu Dekel +2
Multimodal reasoning in vision-language models (VLMs) typically relies on a two-stage process: supervised fine-tuning (SFT) and reinforcement learning (RL). In standard SFT, all to…
Look Where It Matters: High-Resolution Crops Retrieval for Efficient VLMs
Nimrod Shabtay, Moshe Kimhi, Artem Spector +5
Vision-language models (VLMs) typically process images at a native high-resolution, forcing a trade-off between accuracy and computational efficiency: high-resolution inputs captur…
Advancing Speech Understanding in Speech-Aware Language Models with GRPO
Avishai Elmakies, Hagai Aronowitz, Nimrod Shabtay +3
In this paper, we introduce a Group Relative Policy Optimization (GRPO)-based method for training Speech-Aware Large Language Models (SALLMs) on open-format speech understanding ta…
Spoken question answering for visual queries
Nimrod Shabtay, Zvi Kons, Avihu Dekel +3
Question answering (QA) systems are designed to answer natural language questions. Visual QA (VQA) and Spoken QA (SQA) systems extend the textual QA system to accept visual and spo…
Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence
Granite Vision Team, Leonid Karlinsky, Assaf Arbelle +60
We introduce Granite Vision, a lightweight large language model with vision capabilities, specifically designed to excel in enterprise use cases, particularly in visual document un…
Teaching VLMs to Localize Specific Objects from In-context Examples
Sivan Doveh, Nimrod Shabtay, Wei Lin +9
Vision-Language Models (VLMs) have shown remarkable capabilities across diverse visual tasks, including image recognition, video understanding, and Visual Question Answering (VQA)…