6 papers
TripleSumm: Adaptive Triple-Modality Fusion for Video Summarization
Sumin Kim, Hyemin Jeong, Mingu Kang +3
The exponential growth of video content necessitates effective video summarization to efficiently extract key information from long videos. However, current approaches struggle to…
Sounding Highlights: Dual-Pathway Audio Encoders for Audio-Visual Video Highlight Detection
Seohyun Joo, Yoori Oh
Audio-visual video highlight detection aims to automatically identify the most salient moments in videos by leveraging both visual and auditory cues. However, existing models often…
LipSody: Lip-to-Speech Synthesis with Enhanced Prosody Consistency
Jaejun Lee, Yoori Oh, Kyogu Lee
Lip-to-speech synthesis aims to generate speech audio directly from silent facial video by reconstructing linguistic content from lip movements, providing valuable applications in…
Speaking Without Sound: Multi-speaker Silent Speech Voicing with Facial Inputs Only
Jaejun Lee, Yoori Oh, Kyogu Lee
In this paper, we introduce a novel framework for generating multi-speaker speech without relying on any audible inputs. Our approach leverages silent electromyography (EMG) signal…
Hear Your Face: Face-based voice conversion with F0 estimation
Jaejun Lee, Yoori Oh, Injune Hwang +1
This paper delves into the emerging field of face-based voice conversion, leveraging the unique relationship between an individual's facial features and their vocal characteristics…
Distance Sampling-based Paraphraser Leveraging ChatGPT for Text Data Manipulation
Yoori Oh, Yoseob Han, Kyogu Lee
There has been growing interest in audio-language retrieval research, where the objective is to establish the correlation between audio and text modalities. However, most audio-tex…