3 papers
cs.CV2025
Multi-view Gaze Target Estimation
Qiaomu Miao, Vivek Raju Golani, Jingyi Xu +3
This paper presents a method that utilizes multiple camera views for the gaze target estimation (GTE) task. The approach integrates information from different camera views to impro…
cs.CV2025
Importance-Based Token Merging for Efficient Image and Video Generation
Haoyu Wu, Jingyi Xu, Hieu Le +1
Token merging can effectively accelerate various vision systems by processing groups of similar tokens only once and sharing the results across them. However, existing token groupi…
cs.SD2024
Learning Frame-Wise Emotion Intensity for Audio-Driven Talking-Head Generation
Jingyi Xu, Hieu Le, Zhixin Shu +3
Human emotional expression is inherently dynamic, complex, and fluid, characterized by smooth transitions in intensity throughout verbal communication. However, the modeling of suc…