23 papers
Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models
Pengchao Feng, Chao-Hong Tan, Qian Chen +3
Spoken language models (SLMs) enable natural human-computer interaction, but their reasoning ability still lags behind that of text-based large language models, especially on spoke…
AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation
Huadai Liu, Kaicheng Luo, Wen Wang +4
Unifying speech, sound, and music generation in one model is hindered by tradeoffs between fidelity, end-to-end training, in-context conditioning, and variable-length synthesis tha…
STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation
Huadai Liu, Wen Wang, Kaicheng Luo +3
Continuous Variational Autoencoders (VAEs) serve as the fundamental continuous tokenizer for modern neural audio generation systems, enabling high-fidelity reconstruction while pro…
BareWave: Waveform-Native Flow-Matching Text-to-Speech
Wei Fan, Chao-Hong Tan, Qian Chen +5
Removing intermediate representations and separately trained decoding stages has become an important direction in generative modeling. In text-to-speech, however, high-quality syst…
UniVocal: Unified Speech-Singing Code-Switching Synthesis
Yufei Shi, Qian Chen, Wen Wang +3
We propose UniVocal, a unified framework that implicitly infers vocal modes from text context to pioneer Speech-Singing Code-Switching (SCS) Synthesis - a task where transitions ar…
LEGS: Fine-Tuning Teleop-Free VLAs for Humanoid Loco-manipulation in an Embodied Gaussian Splatting World
Hojune Kim, Timothy Chen, Jiankai Sun +4
Training vision-language-action (VLA) policies for humanoid loco-manipulation is constrained by the high cost and complexity of collecting human teleoperation demonstrations. VLA p…