most citedTracker Meets Night: A Transformer Enhancer for UAV Tracking

88 citations · 136 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CV202412 cited

Towards End-to-End Explainable Facial Action Unit Recognition via Vision-Language Joint Learning

Xuri Ge, Junchen Fu, Fuhai Chen +3

Facial action units (AUs), as defined in the Facial Action Coding System (FACS), have received significant research interest owing to their diverse range of applications in facial…

cs.CV202421 cited

3SHNet: Boosting Image-Sentence Retrieval via Visual Semantic-Spatial Self-Highlighting

Xuri Ge, Songpei Xu, Fuhai Chen +4

In this paper, we propose a novel visual Semantic-Spatial Self-Highlighting Network (termed 3SHNet) for high-precision, high-efficiency and high-generalization image-sentence retri…

cs.CV202411 cited

SAR-RARP50: Segmentation of surgical instrumentation and Action Recognition on Robot-Assisted Radical Prostatectomy Challenge

Dimitrios Psychogyios, Emanuele Colleoni, Beatrice Van Amsterdam +47

Surgical tool segmentation and action recognition are fundamental building blocks in many computer-assisted intervention applications, ranging from surgical skills assessment to de…

cs.CV20241 cited

A Bi-Pyramid Multimodal Fusion Method for the Diagnosis of Bipolar Disorders

Guoxin Wang, Sheng Shi, Shan An +5

Previous research on the diagnosis of Bipolar disorder has mainly focused on resting-state functional magnetic resonance imaging. However, their accuracy can not meet the requireme…

eess.IV20232 cited

Multi-Dimension-Embedding-Aware Modality Fusion Transformer for Psychiatric Disorder Clasification

Guoxin Wang, Xuyang Cao, Shan An +5

Deep learning approaches, together with neuroimaging techniques, play an important role in psychiatric disorders classification. Previous studies on psychiatric disorders diagnosis…

cs.CV20231 cited

Temporal Action Localization with Enhanced Instant Discriminability

Dingfeng Shi, Qiong Cao, Yujie Zhong +4

Temporal action detection (TAD) aims to detect all action boundaries and their corresponding categories in an untrimmed video. The unclear boundaries of actions in videos often res…