16 papers
Luna-TTS Family Technical Report
Feng Yin, Shuai Shi, Junjie Zheng +19
Modern text-to-speech (TTS) is dominated by autoregressive (AR) codec language models, whose left-to-right decoding brings latency that grows with utterance length, error accumulat…
On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation
Changhao Cheng, Wei Wang, Wangyou Zhang +4
Continuous speech representations based on Variational Autoencoders (VAEs) have emerged as a promising alternative to traditional spectrogram or discrete token based features for s…
Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation
Inferix Team, Tianyu Feng, Yizeng Han +13
World models serve as core simulators for fields such as agentic AI, embodied AI, and gaming, capable of generating long, physically realistic, and interactive high-quality videos.…
LegoDiffusion: Micro-Serving Text-to-Image Diffusion Workflows
Lingyun Yang, Suyi Li, Tianyu Feng +10
Text-to-image generation executes a diffusion workflow comprising multiple models centered on a base diffusion model. Existing serving systems treat each workflow as an opaque mono…
Exploring Motion-Language Alignment for Text-driven Motion Generation
Ruxi Gu, Zilei Wang, Wei Wang
Text-driven human motion generation aims to synthesize realistic motion sequences that follow textual descriptions. Despite recent advances, accurately aligning motion dynamics wit…
DSFlow: Dual Supervision and Step-Aware Architecture for One-Step Flow Matching Speech Synthesis
Bin Lin, Peng Yang, Chao Yan +5
Flow-matching models have enabled high-quality text-to-speech synthesis, but their iterative sampling process during inference incurs substantial computational cost. Although disti…