13 papers
FlowCTS: On-policy Continuous Trajectory Supervision of Flow Models
Kaiyang Ye, Yuan Ge, Junxiang Zhang +8
While on-policy distillation (OPD) effectively addresses sparse rewards and exposure bias in large language model post-training, its extension to flow models remains underexplored.…
DirectAudioEdit: Inversion-Free Text-Guided Audio Editing via Diffusion Prediction Contrast
Zhengkun Ge, Xiaoqian Liu, Haoran Zhang +5
Text-guided audio editing aims to modify the language-specified acoustic content while preserving edit-irrelevant source components. Existing training-free methods typically rely o…
AvatarForcing: One-Step Streaming Talking Avatars via Local-Future Sliding-Window Denoising
Liyuan Cui, Wentao Hu, Wenyuan Zhang +3
Real-time talking avatar generation requires low latency and minute-level temporal stability. Autoregressive (AR) forcing enables streaming inference but suffers from exposure bias…
On the Emotion Understanding of Synthesized Speech
Yuan Ge, Haishu Zhao, Aokai Hao +10
Emotion is a core paralinguistic feature in voice interaction. It is widely believed that emotion understanding models learn fundamental representations that transfer to synthesize…
When Scaling Fails: Mitigating Audio Perception Decay of LALMs via Multi-Step Perception-Aware Reasoning
Ruixiang Mao, Xiangnan Ma, Dan Chen +12
Test-Time Scaling has shown notable efficacy in addressing complex problems through scaling inference compute. However, within Large Audio-Language Models (LALMs), an unintuitive p…
APR: Penalizing Structural Redundancy in Large Reasoning Models via Anchor-based Process Rewards
Kaiyan Chang, Chenwei Zhu, Yingfeng Luo +7
Test-Time Scaling (TTS) has significantly enhanced the capabilities of Large Reasoning Models (LRMs) but introduces a critical side-effect known as Overthinking. We conduct a preli…