papers

Publications (26)

cs.SD2023

TrimTail: Low-Latency Streaming ASR with Simple but Effective Spectrogram-Level Length Penalty

Xingchen Song, Di Wu, Zhiyong Wu +6

In this paper, we present TrimTail, a simple but effective emission regularization method to improve the latency of streaming ASR models. The core idea of TrimTail is to apply leng…

cs.CL2025

MiMo-VL Technical Report

Core Team, Zihao Yue, Zhenru Lin +71

We open-source MiMo-VL-7B-SFT and MiMo-VL-7B-RL, two powerful vision-language models delivering state-of-the-art performance in both general visual understanding and multimodal rea…

cs.SD2025

Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding

Haoran Zhou, Xingchen Song, Brendan Fahy +9

OpenAI Whisper is a family of robust Automatic Speech Recognition (ASR) models trained on 680,000 hours of audio. However, its encoder-decoder architecture, trained with a sequence…

cs.SD2026

Borderless Long Speech Synthesis

Xingchen Song, Di Wu, Dinghao Zhou +12

Most existing text-to-speech (TTS) systems either synthesize speech sentence by sentence and stitch the results together, or drive synthesis from plain-text dialogues alone. Both a…

cs.CL2026

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis

Xi Wang, Jie Wang, Xingchen Song +8

While generative text-to-speech (TTS) models approach human-level quality, monolithic metrics fail to diagnose fine-grained acoustic artifacts or explain perceptual collapse. To ad…

cs.CL2025

MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining

LLM-Core Xiaomi, :, Bingquan Xia +62

We present MiMo-7B, a large language model born for reasoning tasks, with optimization across both pre-training and post-training stages. During pre-training, we enhance the data p…

eess.AS2024

HydraFormer: One Encoder For All Subsampling Rates

Yaoxun Xu, Xingchen Song, Zhiyong Wu +3

In automatic speech recognition, subsampling is essential for tackling diverse scenarios. However, the inadequacy of a single subsampling rate to address various real-world situati…

cs.CL2024

U2++ MoE: Scaling 4.7x parameters with minimal impact on RTF

Xingchen Song, Di Wu, Binbin Zhang +5

Scale has opened new frontiers in natural language processing, but at a high cost. In response, by learning to only activate a subset of parameters in training and inference, Mixtu…

cs.SD2023

ZeroPrompt: Streaming Acoustic Encoders are Zero-Shot Masked LMs

Xingchen Song, Di Wu, Binbin Zhang +4

In this paper, we present ZeroPrompt (Figure 1-(a)) and the corresponding Prompt-and-Refine strategy (Figure 3), two simple but effective \textbf{training-free} methods to decrease…

cs.SD2021

Non-Autoregressive Transformer ASR with CTC-Enhanced Decoder Input

Xingchen Song, Zhiyong Wu, Yiheng Huang +3

Non-autoregressive (NAR) transformer models have achieved significantly inference speedup but at the cost of inferior accuracy compared to autoregressive (AR) models in automatic s…

cs.SD2023

CB-Conformer: Contextual biasing Conformer for biased word recognition

Yaoxun Xu, Baiji Liu, Qiaochu Huang and +4

Due to the mismatch between the source and target domains, how to better utilize the biased word information to improve the performance of the automatic speech recognition model in…

cs.SD2026

Iterate to Differentiate: Enhancing Discriminability and Reliability in Zero-Shot TTS Evaluation

Shengfan Shen, Di Wu, Xingchen Song +5

Reliable evaluation of modern zero-shot text-to-speech (TTS) models remains challenging. Subjective tests are costly and hard to reproduce, while objective metrics often saturate,…

cs.SD2022

WeNet 2.0: More Productive End-to-End Speech Recognition Toolkit

Binbin Zhang, Di Wu, Zhendong Peng +7

Recently, we made available WeNet, a production-oriented end-to-end speech recognition toolkit, which introduces a unified two-pass (U2) framework and a built-in runtime to address…

cs.SD2022

FusionFormer: Fusing Operations in Transformer for Efficient Streaming Speech Recognition

Xingchen Song, Di Wu, Binbin Zhang +8

The recently proposed Conformer architecture which combines convolution with attention to capture both local and global dependencies has become the \textit{de facto} backbone model…

cs.SD2026

Harness TTS: Towards Context-Aware Expressive Speech Synthesis with Harness Layer

Shengfan Shen, Di Wu, Xingchen Song +5

Expressive speech synthesis for voice assistants requires flexible style control that adapts to explicit requests and broader interaction context. We propose Harness TTS, a lightwe…

eess.AS2023

Spike-Triggered Contextual Biasing for End-to-End Mandarin Speech Recognition

Kaixun Huang, Ao Zhang, Binbin Zhang +3

The attention-based deep contextual biasing method has been demonstrated to effectively improve the recognition performance of end-to-end automatic speech recognition (ASR) systems…

cs.SD2023

LightGrad: Lightweight Diffusion Probabilistic Model for Text-to-Speech

Jie Chen, Xingchen Song, Zhendong Peng +3

Recent advances in neural text-to-speech (TTS) models bring thousands of TTS applications into daily life, where models are deployed in cloud to provide services for customs. Among…

cs.CL2025

MiMo-Audio: Audio Language Models are Few-Shot Learners

Core Team, Dong Zhang, Gang Wang +97

Existing audio language models typically rely on task-specific fine-tuning to accomplish particular audio tasks. In contrast, humans are able to generalize to new audio tasks with…

cs.SD2026

F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation

Dinghao Zhou, Xingchen Song, Di Wu +3

Continuous audio autoencoders reconstruct waveforms well but often produce latents with weak structure for understanding, while self-supervised audio encoders capture semantics but…

cs.SD2022

Fast-U2++: Fast and Accurate End-to-End Speech Recognition in Joint CTC/Attention Frames

Chengdong Liang, Xiao-Lei Zhang, BinBin Zhang +5

Recently, the unified streaming and non-streaming two-pass (U2/U2++) end-to-end model for speech recognition has shown great performance in terms of streaming capability, accuracy…

eess.AS2026

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching

Han Zhu, Wei Kang, Liyong Guo +11

Generating spoken dialogue is inherently more complex than monologue text-to-speech (TTS), as it demands both realistic turn-taking and the maintenance of distinct speaker timbres.…

cs.SD2024

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch

Xingchen Song, Mengtao Xing, Changwei Ma +9

It is well known that LLM-based systems are data-hungry. Recent LLM-based TTS works typically employ complex data processing pipelines to obtain high-quality training data. These s…

cs.CL2020

Speech-XLNet: Unsupervised Acoustic Model Pretraining For Self-Attention Networks

Xingchen Song, Guangsen Wang, Zhiyong Wu +4

Self-attention network (SAN) can benefit significantly from the bi-directional representation learning through unsupervised pretraining paradigms such as BERT and XLNet. In this pa…

eess.AS2026

Qwen-Audio-3.0-Gen-Preview Technical Report

Junyu Dai, Xiaoyue Duan, Xinyue Fan +14

The paper introduces Qwen-Audio-3.0-Gen-Preview, a unified non‑autoregressive model that uses a diffusion transformer and a shared VAE to generate complete mixed‑waveform audio fro…

#audio generation#diffusion models#transformer#variational autoencoder
eess.AS2024

TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch

Xingchen Song, Chengdong Liang, Binbin Zhang +9

Large Automatic Speech Recognition (ASR) models demand a vast number of parameters, copious amounts of data, and significant computational resources during the training process. Ho…

math.MG2023

A new metric associated with the domain boundary

Xingchen Song, Gendi Wang

In this paper, we introduce a new metric which is associated with the domain boundary for a Ptolemy space . Moreover, we study the inclusion relation of the $\ti…