1 paper · 1 filter
Binjie Zhang, Mike Zheng Shou
Tool-augmented vision-language models (VLMs) can solve multimodal, multi-step tasks by calling external tools, yet they remain fragile in practice. Existing works have two common g…