1 paper · 1 filter
Jiyoung Lee, Song Park, Sanghyuk Chun +1
This paper proposes VoxStudio, the first unified and end-to-end speech-to-image model that generates expressive images directly from spoken descriptions by jointly aligning linguis…