2 papers
cs.CV2025
From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Grounded Open-vocabulary Situation Recognition
Chen Cai, Tianyi Liu, Jianjun Gao +5
Recent Multimodal Large Language Models (MLLMs) exhibit strong zero-shot abilities but struggle with complex Grounded Situation Recognition (GSR) and are resource-intensive for edg…
cs.CV2024
CL-HOI: Cross-Level Human-Object Interaction Distillation from Vision Large Language Models
Jianjun Gao, Chen Cai, Ruoyu Wang +4
Human-object interaction (HOI) detection has seen advancements with Vision Language Models (VLMs), but these methods often depend on extensive manual annotations. Vision Large Lang…