1 citations · 4 across the 9 of their papers we have counts for
9 papers
Foley-Omni: A Unified Multimodal Generation Model from Task-Level Audio Synthesis to Complete Video Soundtrack Generation
Ye Tao, Lupeng Liu, Xuenan Xu +6
Recent unified audio generation models can support diverse tasks across speech, sound effects, and music, but most of them still focus on isolated task-level synthesis. However, re…
KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration
Ruicheng Zhang, Kaixi Cong, Jun Zhou +5
Aligning streaming autoregressive (AR) video generators with human preferences is challenging. Existing reinforcement learning methods predominantly rely on noise-based exploration…
Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
Yicheng Ji, Zhizhou Zhong, Jun Zhang +7
Autoregressive (AR) video diffusion models adopt a streaming generation framework, enabling long-horizon video generation with real-time responsiveness, as exemplified by the Self…
AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement
Zhizhou Zhong, Yicheng Ji, Zhe Kong +12
Recently, multi-person video generation has started to gain prominence. While a few preliminary works have explored audio-driven multi-person talking video generation, they often f…
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
Zhuangfei Cheng, Guangyan Zhang, Zehai Tu +6
Foreign accent conversion (FAC) in speech processing remains a challenging task. Building on the remarkable success of large language models (LLMs) in Text-to-Speech (TTS) tasks, t…
Enhancing Segment-Based Speech Emotion Recognition by Deep Self-Learning
Shuiyang Mao, P. C. Ching, Tan Lee
Despite the widespread utilization of deep neural networks (DNNs) for speech emotion recognition (SER), they are severely restricted due to the paucity of labeled data for training…