4 papers
CMTFormer: Marrying Transformer with Hierarchical Information Interaction for RGB-Event Object Detection
Yu Li, Yuenan Hou, Yingmei Wei +2
Event cameras capture sparse brightness changes with high temporal resolution and high dynamic range, compensating for the deficiencies of the conventional RGB frames. However, pre…
EvoDefense: Co-Evolving Black-Box Defense with Large Language Models
Yu Li, Yuenan Hou, Yingmei Wei +2
Large Language Models (LLMs) remain highly vulnerable to diverse attacks, particularly in black-box settings where the internals of target models are inaccessible. Existing black-b…
MoE3D: Mixture of Experts meets Multi-Modal 3D Understanding
Yu Li, Yuenan Hou, Yingmei Wei +4
Multi-modal 3D understanding is a fundamental task in computer vision. Previous multi-modal fusion methods typically employ a single, dense fusion network, struggling to handle the…
Video-based Generalized Category Discovery via Memory-Guided Consistency-Aware Contrastive Learning
Zhang Jing, Pu Nan, Xie Yu Xiang +5
Generalized Category Discovery (GCD) is an emerging and challenging open-world problem that has garnered increasing attention in recent years. Most existing GCD methods focus on di…