7 papers
StepAudio 2.5 Technical Report
Bin Lin, Bo Zhao, Boyong Wu +98
Unified audio-language modeling has emerged as a prominent trend in modern speech systems, promising to bring the reasoning capabilities of large language models to auditory tasks.…
Sparse Compositional Flow Matching by geometric assembly from motion primitives
Yan Tang, Yuanbo Tang, Tingyu Cao +2
Embodied trajectories, such as the executable motion sequences of robotic manipulators, underwater vehicles, and mobile robots, are a fundamental output of embodied AI. Modern gene…
EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation
Songlin Yang, Haobin Zhong, Ruilin Zhang +23
The rapid evolution of generative video foundation models has propelled the field toward professional-grade cinematic synthesis. To achieve such demanding quality, the community tr…
Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training
Yanru Wu, Jianning Wang, Chongxin Gan +1
Training general-purpose Audio Large Language Models (ALLMs) across diverse datasets is essential for holistic audio understanding, yet it faces significant challenges due to datas…
One Model to Translate Them All: Universal Any-to-Any Translation for Heterogeneous Collaborative Perception
Yang Li, Weize Li, Quan Yuan +7
By sharing intermediate features, collaborative perception extends each agent's sensing beyond standalone limits, but real-world feature modality heterogeneity remains a key barrie…
Stabilizing, Scaling & Enhancing MeanFlow for Large-scale Diffusion Distillation
Xiao He, Yang Li, Peizhen Zhang +3
Diffusion models exhibit remarkable generative capability, but their high latency limits practical deployment. Many studies have attempted to reduce sampling steps to accelerate in…