3 papers
cs.LG2025
FANoise: Singular Value-Adaptive Noise Modulation for Robust Multimodal Representation Learning
Jiaoyang Li, Jun Fang, Tianhao Gao +5
Representation learning is fundamental to modern machine learning, powering applications such as text retrieval and multimodal understanding. However, learning robust and generaliz…
cs.SD2025
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement
Xueyao Zhang, Xiaohui Zhang, Kainan Peng +10
The imitation of voice, targeted on specific speech attributes such as timbre and speaking style, is crucial in speech generation. However, existing methods rely heavily on annotat…
cs.SD2024
Efficient Streaming LLM for Speech Recognition
Junteng Jia, Gil Keren, Wei Zhou +6
Recent works have shown that prompting large language models with audio encodings can unlock speech recognition capabilities. However, existing techniques do not scale efficiently,…