activity
20242026
most citedPreference Alignment Improves Language Model-Based TTS

1 citations · 2 across the 10 of their papers we have counts for

collaborators

19 papers

cs.CL2026

Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback

Siddhant Arora, Jinchuan Tian, Jiatong Shi +4

Reinforcement learning from human or AI feedback (RLHF/RLAIF) for speech-in/speech-out dialogue systems (SDS) remains underexplored, with prior work largely limited to single seman…

cs.SD2026

Do Neural Codecs Generalize? A Controlled Study Across Unseen Languages and Non-Speech Tasks

Shih-Heng Wang, Jiatong Shi, Jinchuan Tian +2

This paper investigates three crucial yet underexplored aspects of the generalization capabilities of neural audio codecs (NACs): (i) whether NACs can generalize to unseen language…

cs.SD2026

IKFST: IOO and KOO Algorithms for Accelerated and Precise WFST-based End-to-End Automatic Speech Recognition

Zhuoran Zhuang, Ye Chen, Chao Luo +5

End-to-end automatic speech recognition has become the dominant paradigm in both academia and industry. To enhance recognition performance, the Weighted Finite-State Transducer (WF…

cs.SD2025

PURE Codec: Progressive Unfolding of Residual Entropy for Speech Codec Learning

Jiatong Shi, Haoran Wang, William Chen +4

Neural speech codecs have achieved strong performance in low-bitrate compression, but residual vector quantization (RVQ) often suffers from unstable training and ineffective decomp…

eess.AS2025

BSCodec: A Band-Split Neural Codec for High-Quality Universal Audio Reconstruction

Haoran Wang, Jiatong Shi, Jinchuan Tian +3

Neural audio codecs have recently enabled high-fidelity reconstruction at high compression rates, especially for speech. However, speech and non-speech audio exhibit fundamentally…

cs.SD2025

Speech-DRAME: A Framework for Human-Aligned Benchmarks in Speech Role-Play

Jiatong Shi, Jionghao Han, Yichen Lu +6

Role-play has become a key testbed for generative models, expanding from text-only dialogue to multimodal interaction. Extending role-play to speech captures prosody, emotion, and…