1 citations · 1 across the 4 of their papers we have counts for
Showing eess.ASShow all
3 papers · 1 filter
eess.AS2024
Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
Hanzhao Li, Liumeng Xue, Haohan Guo +6
The multi-codebook speech codec enables the application of large language models (LLM) in TTS but bottlenecks efficiency and robustness due to multi-sequence prediction. To avoid t…
eess.AS2024
AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
Qian Yang, Jin Xu, Wenrui Liu +8
Recently, instruction-following audio-language models have received broad attention for human-audio interaction. However, the absence of benchmarks capable of evaluating audio-cent…
eess.AS2024
SELM: Speech Enhancement Using Discrete Tokens and Language Models
Ziqian Wang, Xinfa Zhu, Zihan Zhang +4
Language models (LMs) have shown superior performances in various speech generation tasks recently, demonstrating their powerful ability for semantic context modeling. Given the in…