1 paper · 1 filter
Gabriel Sarch, Linrong Cai, Qunzhong Wang +3
What does it take to build a visual reasoner that works across charts, science, spatial understanding, and open-ended tasks? The strongest vision-language models (VLMs) suggest tha…