3 papers
cs.CV2025
Audio-Guided Visual Editing with Complex Multi-Modal Prompts
Hyeonyu Kim, Seokhoon Jeong, Seonghee Han +2
Visual editing with diffusion models has made significant progress but often struggles with complex scenarios that textual guidance alone could not adequately describe, highlightin…
cs.CV2023
Generating Realistic Images from In-the-wild Sounds
Taegyeong Lee, Jeonghun Kang, Hyeonyu Kim +1
Representing wild sounds as images is an important but challenging task due to the lack of paired datasets between sound and images and the significant differences in the character…
cs.CV2022
Technical Report for CVPR 2022 LOVEU AQTC Challenge
Hyeonyu Kim, Jongeun Kim, Jeonghun Kang +3
This technical report presents the 2nd winning model for AQTC, a task newly introduced in CVPR 2022 LOng-form VidEo Understanding (LOVEU) challenges. This challenge faces difficult…