3 papers
cs.CV2025
TAVID: Text-Driven Audio-Visual Interactive Dialogue Generation
Ji-Hoon Kim, Junseok Ahn, Doyeop Kwak +2
The objective of this paper is to jointly synthesize interactive videos and conversational speech from text and reference images. With the ultimate goal of building human-like conv…
cs.CV2025
Deep Understanding of Sign Language for Sign to Subtitle Alignment
Youngjoon Jang, Jeongsoo Choi, Junseok Ahn +1
The objective of this work is to align asynchronous subtitles in sign language videos with limited labelled data. To achieve this goal, we propose a novel framework with the follow…
eess.AS2024
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
Jaemin Jung, Junseok Ahn, Chaeyoung Jung +3
We present VoiceDiT, a multi-modal generative model for producing environment-aware speech and audio from text and visual prompts. While aligning speech with text is crucial for in…