activity
20242026
collaborators

13 papers

cs.LG2026

EviDep: Trustworthy Multimodal Depression Estimation via Disentangled Evidential Learning

Fangyuan Liu, Sirui Zhao, Zeyu Zhang +6

Automated multimodal depression estimation in unconstrained environments is inherently challenged by naturalistic noise and complex behavioral variability. Prevailing deterministic…

cs.CV2026

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length

Yubo Huang, Hailong Guo, Fangtai Wu +9

Audio-driven avatar interaction demands real-time, streaming, and infinite-length generation -- capabilities fundamentally at odds with the sequential denoising and long-horizon dr…

cs.CV2026

Tango: Taming Visual Signals for Efficient Video Large Language Models

Shukang Yin, Sirui Zhao, Hanchao Wang +4

Token pruning has emerged as a mainstream approach for developing efficient Video Large Language Models (Video LLMs). This work revisits and advances the two predominant token-prun…

cs.CV2026

ActFER: Agentic Facial Expression Recognition via Active Tool-Augmented Visual Reasoning

Shifeng Liu, Zhengye Zhang, Sirui Zhao +7

Recent advances in Multimodal Large Language Models (MLLMs) have created new opportunities for facial expression recognition (FER), moving it beyond pure label prediction toward re…

cs.SD2025

DMP-TTS: Disentangled multi-modal Prompting for Controllable Text-to-Speech with Chained Guidance

Kang Yin, Chunyu Qiang, Sirui Zhao +6

Controllable text-to-speech (TTS) systems face significant challenges in achieving independent manipulation of speaker timbre and speaking style, often suffering from entanglement…

cs.CV2025

MELLM: A Flow-Guided Large Language Model for Micro-Expression Understanding

Sirui Zhao, Zhengye Zhang, Shifeng Liu +5

Micro-expressions (MEs), brief and low-intensity facial movements revealing concealed emotions, are crucial for affective computing. Despite notable progress in ME recognition, exi…