7 citations · 9 across the 19 of their papers we have counts for
1 paper · 1 filter
Mingjia Shi, Yinhan He, Yaochen Zhu +1
Vision-language models (VLMs) aim to reason by jointly leveraging visual and textual modalities. While allocating additional inference-time computation has proven effective for lar…