activity
20242026
collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2026

AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation

Huadai Liu, Kaicheng Luo, Wen Wang +4

Unifying speech, sound, and music generation in one model is hindered by tradeoffs between fidelity, end-to-end training, in-context conditioning, and variable-length synthesis tha…

eess.AS2026

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation

Huadai Liu, Wen Wang, Kaicheng Luo +3

Continuous Variational Autoencoders (VAEs) serve as the fundamental continuous tokenizer for modern neural audio generation systems, enabling high-fidelity reconstruction while pro…

eess.AS2025

ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing

Huadai Liu, Kaicheng Luo, Jialei Wang +4

While end-to-end video-to-audio generation has greatly improved, producing high-fidelity audio that authentically captures the nuances of visual content remains challenging. Like p…

eess.AS2025

ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control

Shengpeng Ji, Qian Chen, Wen Wang +8

In this paper, we present ControlSpeech, a text-to-speech (TTS) system capable of fully cloning the speaker's voice and enabling arbitrary control and adjustment of speaking style.…

eess.AS2025

OmniAudio: Generating Spatial Audio from 360-Degree Video

Huadai Liu, Tianyi Luo, Kaicheng Luo +11

Traditional video-to-audio generation techniques primarily focus on perspective video and non-spatial audio, often missing the spatial cues necessary for accurately representing so…

eess.AS2025

UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook

Yidi Jiang, Qian Chen, Shengpeng Ji +6

The emergence of audio language models is empowered by neural audio codecs, which establish critical mappings between continuous waveforms and discrete tokens compatible with langu…