most citedQwen2.5-Omni Technical Report

12 citations · 15 across the 5 of their papers we have counts for

collaborators

6 papers

cs.SD2026

LLM-ForcedAligner: A Non-Autoregressive and Accurate LLM-Based Forced Aligner for Multilingual and Long-Form Speech

Bingshen Mu, Xian Shi, Xiong Wang +3

Forced alignment (FA) predicts start and end timestamps for words or characters in speech, but existing methods are language-specific and prone to cumulative temporal shifts. The m…

cs.SD20261 cited

Qwen3-TTS Technical Report

Hangrui Hu, Xinfa Zhu, Ting He +13

In this report, we present the Qwen3-TTS series, a family of advanced multilingual, controllable, robust, and streaming text-to-speech models. Qwen3-TTS supports state-of-the-art 3…

cs.CL20252 cited

Qwen3-Omni Technical Report

Jin Xu, Zhifang Guo, Hangrui Hu +35

We present Qwen3-Omni, a single multimodal model that, for the first time, maintains state-of-the-art performance across text, image, audio, and video without any degradation relat…

eess.AS2025

ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark

He Wang, Linhan Ma, Dake Guo +4

Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluati…

cs.CL202512 cited

Qwen2.5-Omni Technical Report

Jin Xu, Zhifang Guo, Jinzheng He +11

In this report, we present Qwen2.5-Omni, an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously gene…

cs.SD2025

InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training

Dingdong Wang, Jin Xu, Ruihang Chu +6

Recent advancements in speech large language models (SpeechLLMs) have attracted considerable attention. Nonetheless, current methods exhibit suboptimal performance in adhering to s…