1 paper · 1 filter
Muhammad Kamran Janjua, Hugo Silva, Di Niu +1
Multimodal language models (MLLMs) are increasingly paired with vision tools (e.g., depth, flow, correspondence) to enhance visual reasoning. However, despite access to these tool-…