5 papers
ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions
Thomas Thebaud, Junhyeok Lee, Laureano Moro-Velazquez +2
Speaker embeddings, or x-vectors, are widely used to represent speaker identity and speaker-related attributes, but existing embedding extractors are typically descriptive rather t…
Learned Image Compression for Vision-Language-Action Models
Hyeonjun Kim, Jegwang Ryu, Sangbeom Ha +4
Vision-language-action (VLA) models increasingly rely on high-frequency multi-camera observations, making visual communication a major bottleneck for real-time robotic control in b…
MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances
Junhyeok Lee, Helin Wang, Yaohan Guan +4
We introduce MaskVCT, a zero-shot voice conversion (VC) model that offers multi-factor controllability through multiple classifier-free guidances (CFGs). While previous VC models r…
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
Helin Wang, Jiarui Hai, Dading Chong +11
Recent advancements in generative artificial intelligence have significantly transformed the field of style-captioned text-to-speech synthesis (CapTTS). However, adapting CapTTS to…
Improving Test-Time Performance of RVQ-based Neural Codecs
Hyeongju Kim, Junhyeok Lee, Jacob Morton +2
The residual vector quantization (RVQ) technique plays a central role in recent advances in neural audio codecs. These models effectively synthesize high-fidelity audio from a limi…