12 papers
Cross-modal learning for plankton recognition
Joona Kareinen, Veikka Immonen, Tuomas Eerola +5
This paper considers self-supervised cross-modal coordination as a strategy enabling utilization of multiple modalities and large volumes of unlabeled plankton data to build models…
Insights from Visual Cognition: Understanding Human Action Dynamics with Overall Glance and Refined Gaze Transformer
Bohao Xing, Deng Li, Rong Gao +2
Recently, Transformer has made significant progress in various vision tasks. To balance computation and efficiency in video tasks, recent works heavily rely on factorized or window…
DEEMO: De-identity Multimodal Emotion Recognition and Reasoning
Deng Li, Bohao Xing, Xin Liu +3
Emotion understanding is a critical yet challenging task. Most existing approaches rely heavily on identity-sensitive information, such as facial expressions and speech, which rais…
MSF-Mamba: Motion-aware State Fusion Mamba for Efficient Micro-Gesture Recognition
Deng Li, Jun Shao, Bohao Xing +4
Micro-gesture recognition (MGR) targets the identification of subtle and fine-grained human motions and requires accurate modeling of both long-range and local spatiotemporal depen…
Unsupervised Pelage Pattern Unwrapping for Animal Re-identification
Aleksandr Algasov, Ekaterina Nepovinnykh, Fedor Zolotarev +4
Existing individual re-identification methods often struggle with the deformable nature of animal fur or skin patterns which undergo geometric distortions due to body movement and…
EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models
Bohao Xing, Xin Liu, Guoying Zhao +3
Emotion understanding is a critical yet challenging task. Recent advances in Multimodal Large Language Models (MLLMs) have significantly enhanced their capabilities in this area. H…