5 papers
EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization
Lihuang Fang, Yuchen Zou, kebin Jin +2
Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognition (MER) tasks and lifted MER to a new level that is complex emotion un…
Prototype-Regularized Federated Learning for Cross-Domain Aspect Sentiment Triplet Extraction
Zongming Cai, Jianhang Tang, Zhenyong Zhang +3
Aspect Sentiment Triplet Extraction (ASTE) aims to extract all sentiment triplets of aspect terms, opinion terms, and sentiment polarities from a sentence. Existing methods are typ…
OSA: Echocardiography Video Segmentation via Orthogonalized State Update and Anatomical Prior-aware Feature Enhancement
Rui Wang, Huisi Wu, Jing Qin
Accurate and temporally consistent segmentation of the left ventricle from echocardiography videos is essential for estimating the ejection fraction and assessing cardiac function.…
PR-CapsNet: Pseudo-Riemannian Capsule Network with Adaptive Curvature Routing for Graph Learning
Ye Qin, Jingchao Wang, Yang Shi +5
Capsule Networks (CapsNets) show exceptional graph representation capacity via dynamic routing and vectorized hierarchical representations, but they model the complex geometries of…
Learning Semantic-Aware Representation in Visual-Language Models for Multi-Label Recognition with Partial Labels
Haoxian Ruan, Zhihua Xu, Zhijing Yang +3
Multi-label recognition with partial labels (MLR-PL), in which only some labels are known while others are unknown for each image, is a practical task in computer vision, since col…