2 papers
cs.CV2026
HOI-R1: Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection
Junwen Chen, Peilin Xiong, Keiji Yanai
Recent human-object interaction detection (HOID) methods highly require prior knowledge from vision-language models (VLMs) to enhance the interaction recognition capabilities. The…
cs.CV2024
Focusing on what to decode and what to train: SOV Decoding with Specific Target Guided DeNoising and Vision Language Advisor
Junwen Chen, Yingcheng Wang, Keiji Yanai
Recent transformer-based methods achieve notable gains in the Human-object Interaction Detection (HOID) task by leveraging the detection of DETR and the prior knowledge of Vision-L…