papers

Publications (5)

cs.GR2025

FLAP: Fully-controllable Audio-driven Portrait Video Generation through 3D head conditioned diffusion model

Lingzhou Mu, Baiji Liu, Ruonan Zhang +4

Diffusion-based video generation techniques have significantly improved zero-shot talking-head avatar generation, enhancing the naturalness of both head motion and facial expressio…

eess.AS2026

HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech

Sihang Nie, Xiaofen Xing, Rui Xing +5

Recently, Large Language Model (LLM)-based Text-to-Speech (TTS) models have achieved remarkable naturalness. However, the standard Supervised Fine-Tuning paradigm often converges t…

cs.CL2020

Multi-head Monotonic Chunkwise Attention For Online Speech Recognition

Baiji Liu, Songjun Cao, Sining Sun +2

The attention mechanism of the Listen, Attend and Spell (LAS) model requires the whole input sequence to calculate the attention context and thus is not suitable for online speech…

eess.AS2026

HD-PPT: Hierarchical Decoding of Content- and Prompt-Preference Tokens for Instruction-based TTS

Sihang Nie, Xiaofen Xing, Jingyuan Xing +2

Large Language Model (LLM)-based Text-to-Speech (TTS) models have already reached a high degree of naturalness. However, the precision control of TTS inference is still challenging…

cs.SD2023

CB-Conformer: Contextual biasing Conformer for biased word recognition

Yaoxun Xu, Baiji Liu, Qiaochu Huang and +4

Due to the mismatch between the source and target domains, how to better utilize the biased word information to improve the performance of the automatic speech recognition model in…