2 papers
cs.CV2024
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images
Yuchen Yang, Haoran Yan, Yanhao Chen +2
Vision Question Answering (VQA) tasks use images to convey critical information to answer text-based questions, which is one of the most common forms of question answering in real-…
cs.CV2024
Self-Supervised Learning of Deviation in Latent Representation for Co-speech Gesture Video Generation
Huan Yang, Jiahui Chen, Chaofan Ding +5
Gestures are pivotal in enhancing co-speech communication. While recent works have mostly focused on point-level motion transformation or fully supervised motion representations th…