Showing cs.CVShow all
3 papers · 1 filter
cs.CV2024
User-in-the-loop Evaluation of Multimodal LLMs for Activity Assistance
Mrinal Verghese, Brian Chen, Hamid Eghbalzadeh +2
Our research investigates the capability of modern multimodal reasoning models, powered by Large Language Models (LLMs), to facilitate vision-powered assistants for multi-step dail…
cs.CV2023
EgoAdapt: A multi-stream evaluation study of adaptation to real-world egocentric user video
Matthias De Lange, Hamid Eghbalzadeh, Reuben Tan +3
In egocentric action recognition a single population model is typically trained and subsequently embodied on a head-mounted device, such as an augmented reality headset. While this…
cs.CV2023
Pretrained Language Models as Visual Planners for Human Assistance
Dhruvesh Patel, Hamid Eghbalzadeh, Nitin Kamra +3
In our pursuit of advancing multi-modal AI assistants capable of guiding users to achieve complex multi-step goals, we propose the task of "Visual Planning for Assistance (VPA)". G…