3 papers
cs.CV2025
Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive Reasoning
Yian Li, Wentao Tian, Yang Jiao +5
Recently, Multimodal Large Language Models (MLLMs) have achieved significant success across multiple disciplines due to their exceptional instruction-following capabilities and ext…
cs.CV2024
Domain Expansion and Boundary Growth for Open-Set Single-Source Domain Generalization
Pengkun Jiao, Na Zhao, Jingjing Chen +1
Open-set single-source domain generalization aims to use a single-source domain to learn a robust model that can be generalized to unknown target domains with both domain shifts an…
cs.CV2024
Unlocking Textual and Visual Wisdom: Open-Vocabulary 3D Object Detection Enhanced by Comprehensive Guidance from Text and Image
Pengkun Jiao, Na Zhao, Jingjing Chen +1
Open-vocabulary 3D object detection (OV-3DDet) aims to localize and recognize both seen and previously unseen object categories within any new 3D scene. While language and vision f…