activity
20242026
collaborators

5 papers

cs.CV2026

Embedding Rotation Invariance for Provable Multi-Oriented Scene Text Recognition

Zhibin Ma, Pengwen Dai, Yi Liu +3

Multi-oriented text is ubiquitous in real-world scenes and remains a major challenge for scene text recognition (STR). Existing rotation-aware methods explicitly estimate text orie…

cs.CV2026

Beyond Symmetric Fusion: Exploiting Task-Dependent Modality Strengths for RGB-Event Small Object Detection

Ziheng Wang, Chaolang Li, Yutong Yang +5

State-of-the-art RGB-Event detectors improve the detection of small, fast-moving objects by combining complementary features from RGB and Event data, yet they typically fuse the tw…

cs.CV2026

EagleNet: Energy-Aware Fine-Grained Relationship Learning Network for Text-Video Retrieval

Yuhan Chen, Pengwen Dai, Chuan Wang +2

Text-video retrieval tasks have seen significant improvements due to the recent development of large-scale vision-language pre-trained models. Traditional methods primarily focus o…

cs.LG2025

Decoupled Graph Energy-based Model for Node Out-of-Distribution Detection on Heterophilic Graphs

Yuhan Chen, Yihong Luo, Yifan Song +3

Despite extensive research efforts focused on OOD detection on images, OOD detection on nodes in graph learning remains underexplored. The dependence among graph nodes hinders the…

cs.CR2024

Efficient Backdoor Defense in Multimodal Contrastive Learning: A Token-Level Unlearning Method for Mitigating Threats

Kuanrong Liu, Siyuan Liang, Jiawei Liang +2

Multimodal contrastive learning uses various data modalities to create high-quality features, but its reliance on extensive data sources on the Internet makes it vulnerable to back…