collaborators

6 papers

cs.CV2026

Synergizing Understanding and Generation with Interleaved Analyzing-Drafting Thinking

Shengqiong Wu, Bobo Li, Xinkai Wang +6

Unified Vision-Language Models (UVLMs) aim to advance multimodal learning by supporting both understanding and generation within a single framework. However, existing approaches la…

cs.CV2025

VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language Models

Haidong Xu, Guangwei Xu, Zhedong Zheng +7

This paper introduces VimoRAG, a novel video-based retrieval-augmented motion generation framework for motion large language models (LLMs). As motion LLMs face severe out-of-domain…

cs.CV2025

Synergistic Dual Spatial-aware Generation of Image-to-Text and Text-to-Image

Yu Zhao, Hao Fei, Xiangtai Li +6

In the visual spatial understanding (VSU) area, spatial image-to-text (SI2T) and spatial text-to-image (ST2I) are two fundamental tasks that appear in dual form. Existing methods f…

cs.CL2025

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models

Shilin Xu, Yanwei Li, Rui Yang +9

Recent works on large language models (LLMs) have successfully demonstrated the emergence of reasoning capabilities via reinforcement learning (RL). Although recent efforts leverag…

cs.CV2025

On Path to Multimodal Generalist: General-Level and General-Bench

Hao Fei, Yuan Zhou, Juncheng Li +29

The Multimodal Large Language Model (MLLM) is currently experiencing rapid growth, driven by the advanced capabilities of LLMs. Unlike earlier specialists, existing MLLMs are evolv…

cs.CV2025

HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing

Jinbin Bai, Wei Chow, Ling Yang +4

We present HumanEdit, a high-quality, human-rewarded dataset specifically designed for instruction-guided image editing, enabling precise and diverse image manipulations through op…