1 paper · 1 filter
Jaemin Jung, Junseok Ahn, Chaeyoung Jung +3
We present VoiceDiT, a multi-modal generative model for producing environment-aware speech and audio from text and visual prompts. While aligning speech with text is crucial for in…