activity
20242026
collaborators

24 papers

cs.CL2026

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models

Pengchao Feng, Chao-Hong Tan, Qian Chen +3

Spoken language models (SLMs) enable natural human-computer interaction, but their reasoning ability still lags behind that of text-based large language models, especially on spoke…

eess.AS2026

AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation

Huadai Liu, Kaicheng Luo, Wen Wang +4

Unifying speech, sound, and music generation in one model is hindered by tradeoffs between fidelity, end-to-end training, in-context conditioning, and variable-length synthesis tha…

eess.AS2026

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation

Huadai Liu, Wen Wang, Kaicheng Luo +3

Continuous Variational Autoencoders (VAEs) serve as the fundamental continuous tokenizer for modern neural audio generation systems, enabling high-fidelity reconstruction while pro…

eess.AS2026

BareWave: Waveform-Native Flow-Matching Text-to-Speech

Wei Fan, Chao-Hong Tan, Qian Chen +5

Removing intermediate representations and separately trained decoding stages has become an important direction in generative modeling. In text-to-speech, however, high-quality syst…

cs.SD2026

UniVocal: Unified Speech-Singing Code-Switching Synthesis

Yufei Shi, Qian Chen, Wen Wang +3

We propose UniVocal, a unified framework that implicitly infers vocal modes from text context to pioneer Speech-Singing Code-Switching (SCS) Synthesis - a task where transitions ar…

cs.RO2026

LEGS: Fine-Tuning Teleop-Free VLAs for Humanoid Loco-manipulation in an Embodied Gaussian Splatting World

Hojune Kim, Timothy Chen, Jiankai Sun +4

Training vision-language-action (VLA) policies for humanoid loco-manipulation is constrained by the high cost and complexity of collecting human teleoperation demonstrations. VLA p…