9 papers
U-Lens: Supporting User Uncertainty Management in Long-Form LLM Responses
Yu Mei, Qingyue Zhuang, Jie Cai +5
Large language models (LLMs) are increasingly used to generate long-form answers for knowledge-intensive tasks, but users often struggle to decide which parts of a response deserve…
PAGE: Towards Practical Human-level Gaze Target Estimation
Zhoutong Ye, Chengwen Zhang, Zhaibin Cui +10
Gaze target estimation, the task of predicting where a person is looking in a scene, is crucial to understanding human attention and intent. It is a challenging task that combines…
AA: A Multi-view Multimodal Dataset for Screen-based Gaze Estimation
Chang Liu, Jiaqi Liu, Zhoutong Ye +3
We present AA, a multi-view multimodal dataset for screen-based gaze estimation. The dataset captures synchronized facial observations from eight fixed screen-mounted cameras and t…
PAPEL: A Collaborative System for Parental Guidance during Preschool Play-Based English Learning
Xutong Wang, Yu Mei, Qinwei Li +7
Play-based parent-child interaction offers preschoolers rich opportunities for everyday foreign language learning, yet many parents struggle to turn open-ended play into effective…
FLUID: From Ephemeral IDs to Multimodal Semantic Codes for Industrial-Scale Livestreaming Recommendation
Xinhang Yuan, Zexi Huang, Anjia Cao +6
Modern recommender systems rely heavily on ID-based collaborative filtering: each item is represented by a unique ID embedding that accumulates collaborative signals from user inte…
EgoIntrospect: An Egocentric Dataset and Benchmark for User-Centric Internal State Reasoning
Zeyu Wang, Chang Liu, Eduardus Tjitrahardja +22
Despite extensive efforts on egocentric video datasets and benchmarks, understanding users' internal states, which is crucial for enabling seamless AI assistant experiences, remain…