5 papers
Visually-Guided Policy Optimization for Multimodal Reasoning
Zengbin Wang, Feng Xiong, Liang Lin +5
Reinforcement learning with verifiable rewards (RLVR) has significantly advanced the reasoning ability of vision-language models (VLMs). However, the inherent text-dominated nature…
Ace-Skill: Bootstrapping Multimodal Agents with Prioritized and Clustered Evolution
Feng Xiong, Zengbin Wang, Yong Wang +5
Self-evolving agents present a promising path toward continual adaptation by distilling task interactions into reusable knowledge artifacts. In practice, this paradigm remains hind…
AR-MAP: Are Autoregressive Large Language Models Implicit Teachers for Diffusion Large Language Models?
Liang Lin, Feng Xiong, Zengbin Wang +5
Diffusion Large Language Models (DLLMs) have emerged as a powerful alternative to autoregressive models, enabling parallel token generation across multiple positions. However, pref…
Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models
Zengbin Wang, Xuecai Hu, Yong Wang +3
Text-to-image (T2I) models have achieved remarkable success in generating high-fidelity images, but they often fail in handling complex spatial relationships, e.g., spatial percept…
Thinking with Map: Reinforced Parallel Map-Augmented Agent for Geolocalization
Yuxiang Ji, Yong Wang, Ziyu Ma +6
The image geolocalization task aims to predict the location where an image was taken anywhere on Earth using visual clues. Existing large vision-language model (LVLM) approaches le…