5 papers
WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
Shenghao Fu, Yukun Su, Fengyun Rao +3
Open-vocabulary object detection aims to detect arbitrary classes via text prompts. Methods without cross-modal fusion layers (non-fusion) offer faster inference by treating recogn…
LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
Shenghao Fu, Qize Yang, Yuan-Ming Li +3
Long video understanding is still challenging for recent Large Video-Language Models (LVLMs) due to the conflict between long-form temporal understanding and detailed spatial perce…
Reinforcing Video Reasoning with Focused Thinking
Jisheng Dang, Jingze Wu, Teng Wang +6
Recent advancements in reinforcement learning, particularly through Group Relative Policy Optimization (GRPO), have significantly improved multimodal large language models for comp…
A Hierarchical Semantic Distillation Framework for Open-Vocabulary Object Detection
Shenghao Fu, Junkai Yan, Qize Yang +3
Open-vocabulary object detection (OVD) aims to detect objects beyond the training annotations, where detectors are usually aligned to a pre-trained vision-language model, eg, CLIP,…
Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models
Shenghao Fu, Junkai Yan, Qize Yang +3
Recent vision foundation models can extract universal representations and show impressive abilities in various tasks. However, their application on object detection is largely over…