6 papers
SVHighlights: Towards Extremely Long Sport Video Highlight Detection
Donggyu Lee, Youngbin Ki, Jeonghun Kang +1
While highlight detection for long-form videos is of great practical importance, most existing methods remain limited to short-form content, largely due to the absence of a suitabl…
UOTIP: Unbalanced Optimal Transport Map for Unpaired Inverse Problems
Donggyu Lee, Taekyung Lee, Jaewoong Choi
We investigate unpaired image inverse problems, a challenging setting where only independent, non-paired sets of noisy measurements and clean target signals are available for train…
Environmental Understanding Vision-Language Model for Embodied Agent
Jinsik Bang, Jaeyeon Bae, Donggyu Lee +2
Vision-language models (VLMs) have shown strong perception and reasoning abilities for instruction-following embodied agents. However, despite these abilities and their generalizat…
Cross-Modal Emotion Transfer for Emotion Editing in Talking Face Video
Chanhyuk Choi, Taesoo Kim, Donggyu Lee +2
Talking face generation has gained significant attention as a core application of generative models. To enhance the expressiveness and realism of synthesized videos, emotion editin…
PointT2I: LLM-based text-to-image generation via keypoints
Taekyung Lee, Donggyu Lee, Myungjoo Kang
Text-to-image (T2I) generation model has made significant advancements, resulting in high-quality images aligned with an input prompt. However, despite T2I generation's ability to…
Spatial-and-Frequency-aware Restoration method for Images based on Diffusion Models
Kyungsung Lee, Donggyu Lee, Myungjoo Kang
Diffusion models have recently emerged as a promising framework for Image Restoration (IR), owing to their ability to produce high-quality reconstructions and their compatibility w…