5 papers
UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on
Yushe Cao, Shikun Feng, Fei Shen +5
Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics. Dominant approaches cast VVT as mask-cond…
PAGE: Towards Practical Human-level Gaze Target Estimation
Zhoutong Ye, Chengwen Zhang, Zhaibin Cui +10
Gaze target estimation, the task of predicting where a person is looking in a scene, is crucial to understanding human attention and intent. It is a challenging task that combines…
AA: A Multi-view Multimodal Dataset for Screen-based Gaze Estimation
Chang Liu, Jiaqi Liu, Zhoutong Ye +3
We present AA, a multi-view multimodal dataset for screen-based gaze estimation. The dataset captures synchronized facial observations from eight fixed screen-mounted cameras and t…
Customer Service Representative's Perception of the AI Assistant in an Organization's Call Center
Kai Qin, Kexin Du, Yimeng Chen +7
The integration of various AI tools creates a complex socio-technical environment where employee-customer interactions form the core of work practices. This study investigates how…
Prompt2Task: Automating UI Tasks on Smartphones from Textual Prompts
Tian Huang, Chun Yu, Weinan Shi +4
UI task automation enables efficient task execution by simulating human interactions with graphical user interfaces (GUIs), without modifying the existing application code. However…