collaborators

6 papers

cs.AI2026

DiffImaginE: Imagine to Verify Entity Types with Diffusion

Feng Zhang, Feiyu Han, Rongxin Yang +11

Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual evidence. Existing imagine-and…

cs.CV2026

Hierarchical Dual-Change Collaborative Learning for UAV Scene Change Captioning

Fuhai Chen, Pengpeng Huang, Junwen Wu +4

This paper proposes a novel task for UAV scene understanding - UAV Scene Change Captioning (UAV-SCC) - which aims to generate natural language descriptions of semantic changes in d…

cs.MM2024

Multimodal Sentiment Analysis Based on Causal Reasoning

Fuhai Chen, Pengpeng Huang, Xuri Ge +2

With the rapid development of multimedia, the shift from unimodal textual sentiment analysis to multimodal image-text sentiment analysis has obtained academic and industrial attent…

cs.CV2024

Towards End-to-End Explainable Facial Action Unit Recognition via Vision-Language Joint Learning

Xuri Ge, Junchen Fu, Fuhai Chen +3

Facial action units (AUs), as defined in the Facial Action Coding System (FACS), have received significant research interest owing to their diverse range of applications in facial…

cs.CV2024

Hire: Hybrid-modal Interaction with Multiple Relational Enhancements for Image-Text Matching

Xuri Ge, Fuhai Chen, Songpei Xu +3

Image-text matching (ITM) is a fundamental problem in computer vision. The key issue lies in jointly learning the visual and textual representation to estimate their similarity acc…

cs.CV2024

3SHNet: Boosting Image-Sentence Retrieval via Visual Semantic-Spatial Self-Highlighting

Xuri Ge, Songpei Xu, Fuhai Chen +4

In this paper, we propose a novel visual Semantic-Spatial Self-Highlighting Network (termed 3SHNet) for high-precision, high-efficiency and high-generalization image-sentence retri…