5 papers
GUITrans2Act: Understanding User Operational Behaviors from Mobile GUI Interactions with Vision-Language Models
Yudong Zhang, Lei Hu, Daoyang Liu +4
Understanding the digital world on mobile devices is shifting from static UI perception to dynamic action comprehension. This capability enables models to convert visual state tran…
A Text-Native Interface for Generative Video Authoring
Xingyu Bruce Liu, Mira Dontcheva, Dingzeyu Li
Everyone can write their stories in freeform text format -- it's something we all learn in school. Yet storytelling via video requires one to learn specialized and complicated tool…
CoSight: Exploring Viewer Contributions to Online Video Accessibility Through Descriptive Commenting
Ruolin wang, Xingyu Liu, Biao Wang +5
The rapid growth of online video content has outpaced efforts to make visual information accessible to blind and low vision (BLV) audiences. While professional Audio Description (A…
Interacting with Thoughtful AI
Xingyu Bruce Liu, Haijun Xia, Xiang Anthony Chen
We envision the concept of Thoughtful AI, a new human-AI interaction paradigm in which the AI behaves as a continuously thinking entity. Unlike conventional AI systems that operate…
Proactive Conversational Agents with Inner Thoughts
Xingyu Bruce Liu, Shitao Fang, Weiyan Shi +3
One of the long-standing aspirations in conversational AI is to allow them to autonomously take initiatives in conversations, i.e., being proactive. This is especially challenging…