4 papers
SMILE-Next: Teaching Large Language Models to Detect, Classify, and Reason about Laughter
Lee Jung-Mok, Kim Sung-Bin, Joohyun Chang +2
Laughter is a complex social signal that conveys communicative intent beyond amusement. While prior work has focused on isolated laughter analysis tasks, a comprehensive understand…
FacEDiT: Unified Talking Face Editing and Generation via Facial Motion Infilling
Kim Sung-Bin, Joohyun Chang, David Harwath +1
Talking face editing and face generation have often been studied as distinct problems. In this work, we propose viewing both not as separate tasks but as subtasks of a unifying for…
CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition
Jongseo Lee, Joohyun Chang, Dongho Lee +1
We propose Cross-Attention in Audio, Space, and Time (CA^2ST), a transformer-based method for holistic video recognition. Recognizing actions in videos requires both spatial and te…
HERO-VQL: Hierarchical, Egocentric and Robust Visual Query Localization
Joohyun Chang, Soyeon Hong, Hyogun Lee +4
In this work, we tackle the egocentric visual query localization (VQL), where a model should localize the query object in a long-form egocentric video. Frequent and abrupt viewpoin…