works on

From the 1 of 5 linked papers with an AI index.

activity
20242026
collaborators

5 papers

cs.RO2026

RADAR: Closed-Loop Robotic Data Generation via Semantic Planning and Autonomous Causal Environment Reset

Yongzhong Wang, Keyu Zhu, Yong Zhong +3

The paper introduces RADAR, an autonomous closed-loop system that generates large‑scale robot interaction data without human intervention by using vision‑language models for semant…

cs.CV2026

Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception

Yanpeng Sun, Jing Hao, Ke Zhu +6

Training Large Multimodality Models (LMMs) relies on descriptive image caption that connects image and language. Existing methods for generating such captions often rely on distill…

cs.CV2025

On Data Synthesis and Post-training for Visual Abstract Reasoning

Ke Zhu, Yu Wang, Jiangjiang Liu +3

This paper is a pioneering work attempting to address abstract visual reasoning (AVR) problems for large vision-language models (VLMs). We make a common LLaVA-NeXT 7B model capable…

cs.LG2024

Continual SFT Matches Multimodal RLHF with Negative Supervision

Ke Zhu, Yu Wang, Yanpeng Sun +4

Multimodal RLHF usually happens after supervised finetuning (SFT) stage to continually improve vision-language models' (VLMs) comprehension. Conventional wisdom holds its superiori…

cs.CV2024

Self-Supervised Visual Preference Alignment

Ke Zhu, Zheng Ge, Liang Zhao +1

This paper makes the first attempt towards unsupervised preference alignment in Vision-Language Models (VLMs). We generate chosen and rejected responses with regard to the original…