4 papers
UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries
Yijie Zhu, Lingsen Zhang, Zitong Yu +3
Emotional understanding and generation are often treated as separate tasks, yet they are inherently complementary and can mutually enhance each other. In this paper, we propose the…
VLA: Prior-Guided Vision-Language-Action Models via World Knowledge Variation
Yijie Zhu, Jie He, Rui Shao +4
Recent vision-language-action (VLA) models have significantly advanced robotic manipulation by unifying perception, reasoning, and control. To achieve such integration, recent stud…
AdaMHF: Adaptive Multimodal Hierarchical Fusion for Survival Prediction
Shuaiyu Zhang, Xun Lin, Rongxiang Zhang +5
The integration of pathologic images and genomic data for survival analysis has gained increasing attention with advances in multimodal learning. However, current methods often ign…
FusionMamba: Dynamic Feature Enhancement for Multimodal Image Fusion with Mamba
Xinyu Xie, Yawen Cui, Tao Tan +2
Multimodal image fusion aims to integrate information from different imaging techniques to produce a comprehensive, detail-rich single image for downstream vision tasks. Existing m…