7 papers
UniDA3D: A Unified Domain-Adaptive Framework for Multi-View 3D Object Detection
Hongjing Wu, Cheng Chi, Jinlin Wu +3
Camera-only 3D object detection is critical for autonomous driving, offering a cost-effective alternative to LiDAR based methods. In particular, multi-view 3D object detection has…
SA-Person: Text-Based Person Retrieval with Scene-aware Re-ranking
Yingjia Xu, Jinlin Wu, Daming Gao +5
Text-based person retrieval aims to identify a target individual from an image gallery using a natural language description. Existing methods primarily focus on appearance-driven c…
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
Zhenxin Lei, Zhangwei Gao, Changyao Tian +12
Generalist visual captioning goes beyond a simple appearance description task, but requires integrating a series of visual cues into a caption and handling various visual domains.…
PCaM: A Progressive Focus Attention-Based Information Fusion Method for Improving Vision Transformer Domain Adaptation
Zelin Zang, Fei Wang, Liangyu Li +4
Unsupervised Domain Adaptation (UDA) aims to transfer knowledge from a labeled source domain to an unlabeled target domain. Recent UDA methods based on Vision Transformers (ViTs) h…
From Data to Modeling: Fully Open-vocabulary Scene Graph Generation
Zuyao Chen, Jinlin Wu, Zhen Lei +1
We present OvSGTR, a novel transformer-based framework for fully open-vocabulary scene graph generation that overcomes the limitations of traditional closed-set models. Conventiona…
Compile Scene Graphs with Reinforcement Learning
Zuyao Chen, Jinlin Wu, Zhen Lei +2
Next-token prediction is the fundamental principle for training large language models (LLMs), and reinforcement learning (RL) further enhances their reasoning performance. As an ef…