1 paper · 1 filter
Mrinal Verghese, Brian Chen, Hamid Eghbalzadeh +2
Our research investigates the capability of modern multimodal reasoning models, powered by Large Language Models (LLMs), to facilitate vision-powered assistants for multi-step dail…