3 papers
cs.LG2026
Navigating the Emotion Tree: Hierarchical Hyperbolic RAG for Multimodal Emotion Recognition
Zeheng Wang, Bo Zhao, Yijie Zhu +6
Multimodal emotion recognition aims to integrate text, audio, and video sources to understand human affective states. Although multimodal large language models excel at multimodal…
cs.CV2025
MoCA: Identity-Preserving Text-to-Video Generation via Mixture of Cross Attention
Qi Xie, Yongjia Ma, Donglin Di +2
Achieving ID-preserving text-to-video (T2V) generation remains challenging despite recent advances in diffusion-based models. Existing approaches often fail to capture fine-grained…
cs.CV2025
Tuning-Free Long Video Generation via Global-Local Collaborative Diffusion
Yongjia Ma, Junlin Chen, Donglin Di +6
Creating high-fidelity, coherent long videos is a sought-after aspiration. While recent video diffusion models have shown promising potential, they still grapple with spatiotempora…