activity
20242026
collaborators

5 papers

cs.MM2026

Inference-Time Scaling for Joint Audio-Video Generation

Jaemin Jung, Kyeongha Rho, Inkyu Shin +1

Joint audio-video generation aims to synthesize realistic audio-video pairs that are both semantically aligned with text prompts and precisely synchronized. While existing joint au…

eess.AS2025

SCORE: Scaling audio generation using Standardized COmposite REwards

Jaemin Jung, Jaehun Kim, Inkyu Shin +1

The goal of this paper is to enhance Text-to-Audio generation at inference, focusing on generating realistic audio that precisely aligns with text prompts. Despite the rapid advanc…

cs.CV2025

Test-Time Augmentation for Pose-invariant Face Recognition

Jaemin Jung, Youngjoon Jang, Joon Son Chung

The goal of this paper is to enhance face recognition performance by augmenting head poses during the testing phase. Existing methods often rely on training on frontalised images o…

eess.AS2024

VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis

Jaemin Jung, Junseok Ahn, Chaeyoung Jung +3

We present VoiceDiT, a multi-modal generative model for producing environment-aware speech and audio from text and visual prompts. While aligning speech with text is crucial for in…

eess.AS2024

Bridging the Gap between Audio and Text using Parallel-attention for User-defined Keyword Spotting

Youkyum Kim, Jaemin Jung, Jihwan Park +2

This paper proposes a novel user-defined keyword spotting framework that accurately detects audio keywords based on text enrollment. Since audio data possesses additional acoustic…