works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.SD2026

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization

Peijie Chen, Zhuanling Zha, Zhipeng Nie +8

In current zero-shot text-to-speech systems, conventional semantic tokenizers are typically optimized using supervised automatic speech recognition or self-supervised learning obje…

cs.SD2026

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

Weijie Wu, Junbo Li, Lin Li +2

The paper introduces MMAC, a large benchmark of 5,638 audio clips designed to evaluate audio captioning models across multiple capability categories and evaluation dimensions, focu…

cs.SD2026

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations

Peijie Chen, Wenhao Guan, Weijie Wu +7

Zero-shot text-to-speech (TTS) relies on robust speech representations. However, current speech tokenizers face a fundamental trade-off: acoustic codecs preserve high-fidelity audi…

cs.LG2026

Spectral Disentanglement and Enhancement: A Dual-domain Contrastive Framework for Representation Learning

Jinjin Guo, Yexin Li, Zhichao Huang +5

Large-scale multimodal contrastive learning has recently achieved impressive success in learning rich and transferable representations, yet it remains fundamentally limited by the…

cs.LG2025

FANoise: Singular Value-Adaptive Noise Modulation for Robust Multimodal Representation Learning

Jiaoyang Li, Jun Fang, Tianhao Gao +5

Representation learning is fundamental to modern machine learning, powering applications such as text retrieval and multimodal understanding. However, learning robust and generaliz…