5 papers
Environmental Understanding Vision-Language Model for Embodied Agent
Jinsik Bang, Jaeyeon Bae, Donggyu Lee +2
Vision-language models (VLMs) have shown strong perception and reasoning abilities for instruction-following embodied agents. However, despite these abilities and their generalizat…
Cross-Modal Emotion Transfer for Emotion Editing in Talking Face Video
Chanhyuk Choi, Taesoo Kim, Donggyu Lee +2
Talking face generation has gained significant attention as a core application of generative models. To enhance the expressiveness and realism of synthesized videos, emotion editin…
DyaDiT: A Multi-Modal Diffusion Transformer for Socially Favorable Dyadic Gesture Generation
Yichen Peng, Jyun-Ting Song, Siyeol Jung +7
Generating realistic conversational gestures are essential for achieving natural, socially engaging interactions with digital humans. However, existing methods typically map a sing…
Crafting Query-Aware Selective Attention for Single Image Super-Resolution
Junyoung Kim, Youngrok Kim, Siyeol Jung +1
Single Image Super-Resolution (SISR) reconstructs high-resolution images from low-resolution inputs, enhancing image details. While Vision Transformer (ViT)-based models improve SI…
DiffListener: Discrete Diffusion Model for Listener Generation
Siyeol Jung, Taehwan Kim
The listener head generation (LHG) task aims to generate natural nonverbal listener responses based on the speaker's multimodal cues. While prior work either rely on limited modali…