3 citations · 3 across the 4 of their papers we have counts for
4 papers
VideoLLM-online: Online Video Large Language Model for Streaming Video
Joya Chen, Zhaoyang Lv, Shiwei Wu +7
Recent Large Language Models have been enhanced with vision capabilities, enabling them to comprehend images, videos, and interleaved vision-language content. However, the learning…
ASSISTGUI: Task-Oriented Desktop Graphical User Interface Automation
Difei Gao, Lei Ji, Zechen Bai +10
Graphical User Interface (GUI) automation holds significant promise for assisting users with complex tasks, thereby boosting human productivity. Existing works leveraging Large Lan…
AssistQ: Affordance-centric Question-driven Task Completion for Egocentric Assistant
Benita Wong, Joya Chen, You Wu +4
A long-standing goal of intelligent assistants such as AR glasses/robots has been to assist users in affordance-centric real-world scenarios, such as "how can I run the microwave f…
AssistSR: Task-oriented Video Segment Retrieval for Personal AI Assistant
Stan Weixian Lei, Difei Gao, Yuxuan Wang +4
It is still a pipe dream that personal AI assistants on the phone and AR glasses can assist our daily life in addressing our questions like ``how to adjust the date for this watch?…