Publications (42)
PAPEL: A Collaborative System for Parental Guidance during Preschool Play-Based English Learning
Xutong Wang, Yu Mei, Qinwei Li +7
Play-based parent-child interaction offers preschoolers rich opportunities for everyday foreign language learning, yet many parents struggle to turn open-ended play into effective…
Understand Then Memory: A Cognitive Gist-Driven RAG Framework with Global Semantic Diffusion
Pengcheng Zhou, Haochen Li, Zhiqiang Nie +4
Retrieval-Augmented Generation (RAG) effectively mitigates hallucinations in LLMs by incorporating external knowledge. However, the inherent discrete representation of text in exis…
Computing with Smart Rings: A Systematic Literature Review
Zeyu Wang, Ruotong Yu, Xiangyang Wang +13
A smart ring is a wearable electronic device in the form of a ring that incorporates diverse sensors and computing technologies to perform a variety of functions. Designed for use…
Say Your Reason: Extract Contextual Rules In Situ for Context-aware Service Recommendation
Yuxuan Li, Jiahui Li, Lihang Pan +2
This paper introduces SayRea, an interactive system that facilitates the extraction of contextual rules for personalized context-aware service recommendations in mobile scenarios.…
TextOnly: A Unified Function Portal for Text-Related Functions on Smartphones
Minghao Tu, Chun Yu, Xiyuan Shen +3
Text boxes serve as portals to diverse functionalities in today's smartphone applications. However, when it comes to specific functionalities, users always need to navigate through…
G-VOILA: Gaze-Facilitated Information Querying in Daily Scenarios
Zeyu Wang, Yuanchun Shi, Yuntao Wang +6
Modern information querying systems are progressively incorporating multimodal inputs like vision and audio. However, the integration of gaze -- a modality deeply linked to user in…
Critiquing Self-report Practices for Human Mental and Wellbeing Computing at Ubicomp
Nan Gao, Soundariya Ananthan, Chun Yu +2
Computing human mental and wellbeing is crucial to various domains, including health, education, and entertainment. However, the reliance on self-reporting in traditional research…
GestureGPT: Toward Zero-Shot Free-Form Hand Gesture Understanding with Large Language Model Agents
Xin Zeng, Xiaoyu Wang, Tengxiang Zhang +3
Existing gesture interfaces only work with a fixed set of gestures defined either by interface designers or by users themselves, which introduces learning or demonstration efforts…
Adapting AI to the Moment: Understanding the Dynamics of Parent-AI Collaboration Modes in Real-Time Conversations with Children
Yu Mei, Ziyao Zhang, Qingyang Wan +5
Parent-AI collaboration to support real-time conversations with children is challenging due to the sensitivity and open-ended nature of such interactions. Existing systems often si…
Prompt2Task: Automating UI Tasks on Smartphones from Textual Prompts
Tian Huang, Chun Yu, Weinan Shi +4
UI task automation enables efficient task execution by simulating human interactions with graphical user interfaces (GUIs), without modifying the existing application code. However…
SituFont: A Just-in-Time Adaptive Intervention System for Enhancing Mobile Readability in Situational Visual Impairments
Jingruo Chen, Kexin Nie, Mingshan Zhang +5
Situational visual impairments (SVIs) hinder mobile readability, causing discomfort and limiting information access. Building on prior work in adaptive typography and accessibility…
The Homework Wars: Exploring Emotions, Behaviours, and Conflicts in Parent-Child Homework Interactions
Nan Gao, Yibin Liu, Xin Tang +8
Parental involvement in homework is a crucial aspect of family education, but it often triggers emotional strain and conflicts. Despite growing concern over its impact on family we…
PAGE: Towards Practical Human-level Gaze Target Estimation
Zhoutong Ye, Chengwen Zhang, Zhaibin Cui +10
Gaze target estimation, the task of predicting where a person is looking in a scene, is crucial to understanding human attention and intent. It is a challenging task that combines…
PoseAugment: Generative Human Pose Data Augmentation with Physical Plausibility for IMU-based Motion Capture
Zhuojun Li, Chun Yu, Chen Liang +1
The data scarcity problem is a crucial factor that hampers the model performance of IMU-based human motion capture. However, effective data augmentation for IMU-based motion captur…
Robust Linear Regression: A Review and Comparison
Chun Yu, Weixin Yao, Xue Bai
Ordinary least-squares (OLS) estimators for a linear model are very sensitive to unusual values in the design space or outliers among y values. Even one single atypical value may h…
AR Secretary Agent: Real-time Memory Augmentation via LLM-powered Augmented Reality Glasses
Raphaël A. El Haddad, Zeyu Wang, Yeonsu Shin +3
Interacting with a significant number of individuals on a daily basis is commonplace for many professionals, which can lead to challenges in recalling specific details: Who is this…
Pursuing Sources of Heterogeneity in Modeling Clustered Population
Yan Li, Chun Yu, Yize Zhao +3
Researchers often have to deal with heterogeneous population with mixed regression relationships, increasingly so in the era of data explosion. In such problems, when there are man…
SonarWatch: Field sensing technique for smartwatches based on ultrasound and motion
Yingtian Shi, Chun Yu, Xuyang Lu +3
A smartwatch worn continuously on the wrist has the potential to perceive rich interactive gestures and natural behaviors of the user. Unfortunately, the current interaction functi…
MOAT: Evaluating LMMs for Capability Integration and Instruction Grounding
Zhoutong Ye, Mingze Sun, Huan-ang Gao +9
Large multimodal models (LMMs) have demonstrated significant potential as generalists in vision-language (VL) tasks. However, adoption of LMMs in real-world tasks is hindered by th…
TaskSense: Cognitive Chain Modeling and Difficulty Estimation for GUI Tasks
Yiwen Yin, Zhian Hu, Xiaoxi Xu +4
Measuring GUI task difficulty is crucial for user behavior analysis and agent capability evaluation. Yet, existing benchmarks typically quantify difficulty based on motor actions (…
Enabling Voice-Accompanying Hand-to-Face Gesture Recognition with Cross-Device Sensing
Zisu Li, Cheng Liang, Yuntao Wang +5
Gestures performed accompanying the voice are essential for voice interaction to convey complementary semantics for interaction purposes such as wake-up state and input modality. I…
AngleSizer: Enhancing Spatial Scale Perception for the Visually Impaired with an Interactive Smartphone Assistant
Xiaoqing Jing, Chun Yu, Kun Yue +6
Spatial perception, particularly at small and medium scales, is an essential human sense but poses a significant challenge for the blind and visually impaired (BVI). Traditional le…
Modeling the Trade-off of Privacy Preservation and Activity Recognition on Low-Resolution Images
Yuntao Wang, Zirui Cheng, Xin Yi +7
A computer vision system using low-resolution image sensors can provide intelligent services (e.g., activity recognition) but preserve unnecessary visual privacy information from t…
Revamp: Enhancing Accessible Information Seeking Experience of Online Shopping for Blind or Low Vision Users
Ruolin Wang, Zixuan Chen, Mingrui "Ray" Zhang +5
Online shopping has become a valuable modern convenience, but blind or low vision (BLV) users still face significant challenges using it, because of: 1) inadequate image descriptio…
AA: A Multi-view Multimodal Dataset for Screen-based Gaze Estimation
Chang Liu, Jiaqi Liu, Zhoutong Ye +3
We present AA, a multi-view multimodal dataset for screen-based gaze estimation. The dataset captures synchronized facial observations from eight fixed screen-mounted cameras and t…
UbiPhysio: Support Daily Functioning, Fitness, and Rehabilitation with Action Understanding and Feedback in Natural Language
Chongyang Wang, Yuan Feng, Lingxiao Zhong +8
We introduce UbiPhysio, a milestone framework that delivers fine-grained action description and feedback in natural language to support people's daily functioning, fitness, and reh…
LLMartini: Seamless and Interactive Leveraging of Multiple LLMs through Comparison and Composition
Yingtian Shi, Jinda Yang, Yuhan Wang +4
The growing diversity of large language models (LLMs) means users often need to compare and combine outputs from different models to obtain higher-quality or more comprehensive res…
Customer Service Representative's Perception of the AI Assistant in an Organization's Call Center
Kai Qin, Kexin Du, Yimeng Chen +7
The integration of various AI tools creates a complex socio-technical environment where employee-customer interactions form the core of work practices. This study investigates how…
Twitch Third-Party Developers' Support Seeking and Provision Practices on Discord
Jie Cai, He Zhang, Yueyan Liu +2
Third-party developers (TPDs) often turn to online communities for support when they can't get immediate responses from the platform. Twitch, as a leading live streaming platform,…
Bridging the gap between natural user expression with complex automation programming in smart homes
Yingtian Shi, Xiaoyi Liu, Chun Yu +4
A long-standing challenge in end-user programming (EUP) is to trade off between natural user expression and the complexity of programming tasks. As large language models (LLMs) are…
Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation
Yushe Cao, Dianxi Shi, Xing Fu +5
While significant progress has been achieved in multimodal facial generation using semantic masks and textual descriptions, conventional feature fusion approaches often fail to ena…
KeySense: LLM-Powered Hands-Down, Ten-Finger Typing on Commodity Touchscreens
Tony Li, Yan Ma, Zhuojun Li +3
Existing touchscreen software keyboards prevent users from resting their hands, forcing slow and fatiguing index-finger tapping ("chicken typing") instead of familiar hands-down te…
Division of Labor and Collaboration Between Parents in Family Education
Ziyi Wang, Congrong Zhang, Jingying Deng +5
Homework tutoring work is a demanding and often conflict-prone practice in family life, and parents often lack targeted support for managing its cognitive and emotional burdens. Th…
HiSync: Spatio-Temporally Aligning Hand Motion from Wearable IMU and On-Robot Camera for Command Source Identification in Long-Range HRI
Chengwen Zhang, Chun Yu, Borong Zhuang +9
Long-range Human-Robot Interaction (HRI) remains underexplored. Within it, Command Source Identification (CSI) - determining who issued a command - is especially challenging due to…
Leveraging Large Language Models for Generating Mobile Sensing Strategies in Human Behavior Modeling
Nan Gao, Zhuolei Yu, Yue Xu +4
Mobile sensing plays a crucial role in generating digital traces to understand human daily lives. However, studying behaviours like mood or sleep quality in smartphone users requir…
AutoTask: Executing Arbitrary Voice Commands by Exploring and Learning from Mobile GUI
Lihang Pan, Bowen Wang, Chun Yu +3
Voice command interfaces (VCIs) have gained increasing importance, enabling hands-free and eyes-free interaction with digital devices. However, the inherent complexity in construct…
U-Lens: Supporting User Uncertainty Management in Long-Form LLM Responses
Yu Mei, Qingyue Zhuang, Jie Cai +5
Large language models (LLMs) are increasingly used to generate long-form answers for knowledge-intensive tasks, but users often struggle to decide which parts of a response deserve…
PACEE: Parent-Centered AI Scaffolding for Emotion Education in Early Childhood Conversations
Yu Mei, Xutong Wang, Ziyao Zhang +8
Emotion education is critical for children aged 3 to 6. However, existing technologies largely focus on children's direct interaction with AI, overlooking the central role of paren…
UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on
Yushe Cao, Shikun Feng, Fei Shen +5
Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics. Dominant approaches cast VVT as mask-cond…
CasualGaze: Towards Modeling and Recognizing Casual Gaze Behavior for Efficient Gaze-based Object Selection
Yingtian Shi, Yukang Yan, Zisu Li +4
We present CasualGaze, a novel eye-gaze-based target selection technique to support natural and casual eye-gaze input. Unlike existing solutions that require users to keep the eye-…
MindShift: Leveraging Large Language Models for Mental-States-Based Problematic Smartphone Use Intervention
Ruolan Wu, Chun Yu, Xiaole Pan +9
Problematic smartphone use negatively affects physical and mental health. Despite the wide range of prior research, existing persuasive techniques are not flexible enough to provid…
A Human-Computer Collaborative Tool for Training a Single Large Language Model Agent into a Network through Few Examples
Lihang Pan, Yuxuan Li, Chun Yu +1
The capabilities of a single large language model (LLM) agent for solving a complex task are limited. Connecting multiple LLM agents to a network can effectively improve overall pe…