3 papers
eess.AS2026
Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec
Yanzhou Ren, Noboru Harada, Daiki Takeuchi +6
Neural audio codec (NAC) is essential for reconstructing high-quality speech signals and generating discrete representations for downstream speech language models. However, ensurin…
eess.AS2026
TTA: Transcribe, Translate and Alignment for Cross-lingual Speech Representation
Wei Liu, Jiahong Li, Yiwen Shao +1
Speech-LLM models have demonstrated great performance in multi-modal and multi-task speech understanding. A typical speech-LLM paradigm is integrating speech modality with a large…
cs.CL2025
AzeroS: Extending LLM to Speech with Self-Generated Instruction-Free Tuning
Yiwen Shao, Wei Liu, Jiahong Li +4
Extending large language models (LLMs) to the speech domain has recently gained significant attention. A typical approach connects a pretrained LLM with an audio encoder through a…