6 papers
U-Lens: Supporting User Uncertainty Management in Long-Form LLM Responses
Yu Mei, Qingyue Zhuang, Jie Cai +5
Large language models (LLMs) are increasingly used to generate long-form answers for knowledge-intensive tasks, but users often struggle to decide which parts of a response deserve…
PAGE: Towards Practical Human-level Gaze Target Estimation
Zhoutong Ye, Chengwen Zhang, Zhaibin Cui +10
Gaze target estimation, the task of predicting where a person is looking in a scene, is crucial to understanding human attention and intent. It is a challenging task that combines…
AA: A Multi-view Multimodal Dataset for Screen-based Gaze Estimation
Chang Liu, Jiaqi Liu, Zhoutong Ye +3
We present AA, a multi-view multimodal dataset for screen-based gaze estimation. The dataset captures synchronized facial observations from eight fixed screen-mounted cameras and t…
PAPEL: A Collaborative System for Parental Guidance during Preschool Play-Based English Learning
Xutong Wang, Yu Mei, Qinwei Li +7
Play-based parent-child interaction offers preschoolers rich opportunities for everyday foreign language learning, yet many parents struggle to turn open-ended play into effective…
HiSync: Spatio-Temporally Aligning Hand Motion from Wearable IMU and On-Robot Camera for Command Source Identification in Long-Range HRI
Chengwen Zhang, Chun Yu, Borong Zhuang +9
Long-range Human-Robot Interaction (HRI) remains underexplored. Within it, Command Source Identification (CSI) - determining who issued a command - is especially challenging due to…
MOAT: Evaluating LMMs for Capability Integration and Instruction Grounding
Zhoutong Ye, Mingze Sun, Huan-ang Gao +9
Large multimodal models (LMMs) have demonstrated significant potential as generalists in vision-language (VL) tasks. However, adoption of LMMs in real-world tasks is hindered by th…