1 citations · 1 across the 7 of their papers we have counts for
1 paper · 1 filter
Shaked Perek, Ben Wiesel, Avihu Dekel +2
Multimodal reasoning in vision-language models (VLMs) typically relies on a two-stage process: supervised fine-tuning (SFT) and reinforcement learning (RL). In standard SFT, all to…