1 paper · 1 filter
Aritra Dutta, Somak Aditya
Vision-language models (VLMs) increasingly rely on chain-of-thought (CoT) reasoning to solve complex multimodal tasks, but their large parameter sizes make deployment expensive. St…