inference-time refinement 1multimodal reasoning 1reinforcement learning 1self-verification 1vision-language models 1
From the 1 of 5 linked papers with an AI index.
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Inference Compute-Optimal Video Vision Language Models
Peiqi Wang, ShengYun Peng, Xuewen Zhang +5
This work investigates the optimal allocation of inference compute across three key scaling factors in video vision language models: language model size, frame count, and the numbe…
cs.CV2024
CompCap: Improving Multimodal Large Language Models with Composite Captions
Xiaohui Chen, Satya Narayan Shukla, Mahmoud Azab +8
How well can Multimodal Large Language Models (MLLMs) understand composite images? Composite images (CIs) are synthetic visuals created by merging multiple visual elements, such as…