4 papers
EventCrab: Harnessing Frame and Point Synergy for Event-based Action Recognition and Beyond
Meiqi Cao, Xiangbo Shu, Jiachao Zhang +3
Event-based Action Recognition (EAR) possesses the advantages of high-temporal resolution capturing and privacy preservation compared with traditional action recognition. Current l…
Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction Recognition
Shiyu Xuan, Dongkai Wang, Zechao Li +1
Zero-shot Human-object interaction (HOI) detection aims to locate humans and objects in images and recognize their interactions. While advances in open-vocabulary object detection…
Exploring Effective Factors for Improving Visual In-Context Learning
Yanpeng Sun, Qiang Chen, Xiaofan Li +3
The In-Context Learning (ICL) is to understand a new task via a few demonstrations (aka. prompt) and predict new inputs without tuning the models. While it has been widely studied…
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking
Shiyu Xuan, Zechao Li, Jinhui Tang
Multi-modal object tracking integrates auxiliary modalities such as depth, thermal infrared, event flow, and language to provide additional information beyond RGB images, showing g…