collaborators

6 papers

cs.CV2025

RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping

Dongming Wu, Yanping Fu, Saike Huang +8

General robotic grasping systems require accurate object affordance perception in diverse open-world scenarios following human instructions. However, current studies suffer from th…

cs.CV2025

Language Prompt for Autonomous Driving

Dongming Wu, Wencheng Han, Yingfei Liu +4

A new trend in the computer vision community is to capture objects of interest following flexible human command represented by a natural language prompt. However, the progress of u…

cs.RO2025

Glad: A Streaming Scene Generator for Autonomous Driving

Bin Xie, Yingfei Liu, Tiancai Wang +2

The generation and simulation of diverse real-world scenes have significant application value in the field of autonomous driving, especially for the corner cases. Recently, researc…

cs.CV2024

Reconstructive Visual Instruction Tuning

Haochen Wang, Anlin Zheng, Yucheng Zhao +4

This paper introduces reconstructive visual instruction tuning (ROSS), a family of Large Multimodal Models (LMMs) that exploit vision-centric supervision signals. In contrast to co…

cs.CV2024

SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control

Binyuan Huang, Yuqing Wen, Yucheng Zhao +9

Autonomous driving progress relies on large-scale annotated datasets. In this work, we explore the potential of generative models to produce vast quantities of freely-labeled data…

cs.CV2024

Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models

Meng Cao, Yuyang Liu, Yingfei Liu +6

Instruction tuning constitutes a prevalent technique for tailoring Large Vision Language Models (LVLMs) to meet individual task requirements. To date, most of the existing approach…