5 papers
Inference-Time Scaling for Joint Audio-Video Generation
Jaemin Jung, Kyeongha Rho, Inkyu Shin +1
Joint audio-video generation aims to synthesize realistic audio-video pairs that are both semantically aligned with text prompts and precisely synchronized. While existing joint au…
SCORE: Scaling audio generation using Standardized COmposite REwards
Jaemin Jung, Jaehun Kim, Inkyu Shin +1
The goal of this paper is to enhance Text-to-Audio generation at inference, focusing on generating realistic audio that precisely aligns with text prompts. Despite the rapid advanc…
Test-Time Augmentation for Pose-invariant Face Recognition
Jaemin Jung, Youngjoon Jang, Joon Son Chung
The goal of this paper is to enhance face recognition performance by augmenting head poses during the testing phase. Existing methods often rely on training on frontalised images o…
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
Jaemin Jung, Junseok Ahn, Chaeyoung Jung +3
We present VoiceDiT, a multi-modal generative model for producing environment-aware speech and audio from text and visual prompts. While aligning speech with text is crucial for in…
Bridging the Gap between Audio and Text using Parallel-attention for User-defined Keyword Spotting
Youkyum Kim, Jaemin Jung, Jihwan Park +2
This paper proposes a novel user-defined keyword spotting framework that accurately detects audio keywords based on text enrollment. Since audio data possesses additional acoustic…