activity
20202026
most citedEigenEmo: Spectral Utterance Representation Using Dynamic Mode Decomposition for Speech Emotion Classification

1 citations · 4 across the 9 of their papers we have counts for

collaborators

9 papers

cs.SD2026

Foley-Omni: A Unified Multimodal Generation Model from Task-Level Audio Synthesis to Complete Video Soundtrack Generation

Ye Tao, Lupeng Liu, Xuenan Xu +6

Recent unified audio generation models can support diverse tasks across speech, sound effects, and music, but most of them still focus on isolated task-level synthesis. However, re…

cs.CV2026

KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration

Ruicheng Zhang, Kaixi Cong, Jun Zhou +5

Aligning streaming autoregressive (AR) video generators with human preferences is challenging. Existing reinforcement learning methods predominantly rely on noise-based exploration…

cs.CV2026

Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models

Yicheng Ji, Zhizhou Zhong, Jun Zhang +7

Autoregressive (AR) video diffusion models adopt a streaming generation framework, enabling long-horizon video generation with real-time responsiveness, as exemplified by the Self…

cs.CV2025

AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement

Zhizhou Zhong, Yicheng Ji, Zhe Kong +12

Recently, multi-person video generation has started to gain prominence. While a few preliminary works have explored audio-driven multi-person talking video generation, they often f…

eess.AS2025

SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech

Zhuangfei Cheng, Guangyan Zhang, Zehai Tu +6

Foreign accent conversion (FAC) in speech processing remains a challenging task. Building on the remarkable success of large language models (LLMs) in Text-to-Speech (TTS) tasks, t…

eess.AS20211 cited

Enhancing Segment-Based Speech Emotion Recognition by Deep Self-Learning

Shuiyang Mao, P. C. Ching, Tan Lee

Despite the widespread utilization of deep neural networks (DNNs) for speech emotion recognition (SER), they are severely restricted due to the paucity of labeled data for training…