3 papers
cs.SD2026
VividVoice: A Unified Framework for Scene-Aware Visually-Driven Speech Synthesis
Chengyuan Ma, Jiawei Jin, Ruijie Xiong +3
We introduce and define a novel task-Scene-Aware Visually-Driven Speech Synthesis, aimed at addressing the limitations of existing speech generation models in creating immersive au…
cs.CL2025
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
Canxiang Yan, Chunxiang Jin, Dawei Huang +22
Existing speech models suffer from competing requirements on token representations by understanding and generation tasks. This discrepancy in representation prevents speech languag…
cs.AI2025
Ming-Omni: A Unified Multimodal Model for Perception and Generation
Inclusion AI, Biao Gong, Cheng Zou +55
We propose Ming-Omni, a unified multimodal model capable of processing images, text, audio, and video, while demonstrating strong proficiency in both speech and image generation. M…