activity
20242026
collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

UniDA3D: A Unified Domain-Adaptive Framework for Multi-View 3D Object Detection

Hongjing Wu, Cheng Chi, Jinlin Wu +3

Camera-only 3D object detection is critical for autonomous driving, offering a cost-effective alternative to LiDAR based methods. In particular, multi-view 3D object detection has…

cs.CV2025

SA-Person: Text-Based Person Retrieval with Scene-aware Re-ranking

Yingjia Xu, Jinlin Wu, Daming Gao +5

Text-based person retrieval aims to identify a target individual from an image gallery using a natural language description. Existing methods primarily focus on appearance-driven c…

cs.CV2025

MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites

Zhenxin Lei, Zhangwei Gao, Changyao Tian +12

Generalist visual captioning goes beyond a simple appearance description task, but requires integrating a series of visual cues into a caption and handling various visual domains.…

cs.CV2025

From Data to Modeling: Fully Open-vocabulary Scene Graph Generation

Zuyao Chen, Jinlin Wu, Zhen Lei +1

We present OvSGTR, a novel transformer-based framework for fully open-vocabulary scene graph generation that overcomes the limitations of traditional closed-set models. Conventiona…

cs.CV2025

Compile Scene Graphs with Reinforcement Learning

Zuyao Chen, Jinlin Wu, Zhen Lei +2

Next-token prediction is the fundamental principle for training large language models (LLMs), and reinforcement learning (RL) further enhances their reasoning performance. As an ef…

cs.CV2025

What Makes a Scene ? Scene Graph-based Evaluation and Feedback for Controllable Generation

Zuyao Chen, Jinlin Wu, Zhen Lei +1

While text-to-image generation has been extensively studied, generating images from scene graphs remains relatively underexplored, primarily due to challenges in accurately modelin…