most citedMing-Omni: A Unified Multimodal Model for Perception and Generation

1 citations · 1 across the 4 of their papers we have counts for

collaborators

7 papers

cs.SD2026

StepAudio 3 Gen Technical Report

Bin Lin, Bo Zhao, Boyang Wang +68

We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe spee…

eess.AS2026

Robust Speech Emotion Recognition under Tone-Word Conflict: A Benchmark and Framework

Xiaojiang Peng, Dawei Huang, Yongjie Lv +5

Speech emotion recognition (SER) is a crucial component of human-computer interaction, attracting extensive attention from both industry and academia. However, existing SER systems…

cs.SD2026

VividVoice: A Unified Framework for Scene-Aware Visually-Driven Speech Synthesis

Chengyuan Ma, Jiawei Jin, Ruijie Xiong +3

We introduce and define a novel task-Scene-Aware Visually-Driven Speech Synthesis, aimed at addressing the limitations of existing speech generation models in creating immersive au…

cs.SD2026

When Tone and Words Disagree: Towards Robust Speech Emotion Recognition under Acoustic-Semantic Conflict

Dawei Huang, Yongjie Lv, Ruijie Xiong +2

Speech Emotion Recognition (SER) systems often assume congruence between vocal emotion and lexical semantics. However, in real-world interactions, acoustic-semantic conflict is com…

cs.CL2025

Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation

Canxiang Yan, Chunxiang Jin, Dawei Huang +22

Existing speech models suffer from competing requirements on token representations by understanding and generation tasks. This discrepancy in representation prevents speech languag…

cs.CV2025

Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation

Inclusion AI, :, Bowen Ma +73

We propose Ming-Flash-Omni, an upgraded version of Ming-Omni, built upon a sparser Mixture-of-Experts (MoE) variant of Ling-Flash-2.0 with 100 billion total parameters, of which on…