6 papers
DiffImaginE: Imagine to Verify Entity Types with Diffusion
Feng Zhang, Feiyu Han, Rongxin Yang +11
Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence. Existing imagine-and…
Hierarchical Dual-Change Collaborative Learning for UAV Scene Change Captioning
Fuhai Chen, Pengpeng Huang, Junwen Wu +4
This paper proposes a novel task for UAV scene understanding - UAV Scene Change Captioning (UAV-SCC) - which aims to generate natural language descriptions of semantic changes in d…
Multimodal Sentiment Analysis Based on Causal Reasoning
Fuhai Chen, Pengpeng Huang, Xuri Ge +2
With the rapid development of multimedia, the shift from unimodal textual sentiment analysis to multimodal image-text sentiment analysis has obtained academic and industrial attent…
Towards End-to-End Explainable Facial Action Unit Recognition via Vision-Language Joint Learning
Xuri Ge, Junchen Fu, Fuhai Chen +3
Facial action units (AUs), as defined in the Facial Action Coding System (FACS), have received significant research interest owing to their diverse range of applications in facial…
Hire: Hybrid-modal Interaction with Multiple Relational Enhancements for Image-Text Matching
Xuri Ge, Fuhai Chen, Songpei Xu +3
Image-text matching (ITM) is a fundamental problem in computer vision. The key issue lies in jointly learning the visual and textual representation to estimate their similarity acc…
3SHNet: Boosting Image-Sentence Retrieval via Visual Semantic-Spatial Self-Highlighting
Xuri Ge, Songpei Xu, Fuhai Chen +4
In this paper, we propose a novel visual Semantic-Spatial Self-Highlighting Network (termed 3SHNet) for high-precision, high-efficiency and high-generalization image-sentence retri…