most citedVLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

3 citations · 5 across the 4 of their papers we have counts for

collaborators

7 papers

cs.CV2025

VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs

Peng Liu, Haozhan Shen, Chunxin Fang +3

Vision-Language Models (VLMs) excel at high-level scene understanding but falter on fine-grained perception tasks requiring precise localization. This failure stems from a fundamen…

cs.CL2025

Unifying Language Agent Algorithms with Graph-based Orchestration Engine for Reproducible Agent Research

Qianqian Zhang, Jiajia Liao, Heting Ying +9

Language agents powered by large language models (LLMs) have demonstrated remarkable capabilities in understanding, reasoning, and executing complex tasks. However, developing robu…

cs.CV20253 cited

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Haozhan Shen, Peng Liu, Jingcheng Li +9

Recently DeepSeek R1 has shown that reinforcement learning (RL) can substantially improve the reasoning capabilities of Large Language Models (LLMs) through a simple yet effective…

cs.CV2025

GeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote Sensing

Zilun Zhang, Haozhan Shen, Tiancheng Zhao +7

The application of Vision-Language Models (VLMs) in remote sensing (RS) has demonstrated significant potential in traditional tasks such as scene classification, object detection,…

cs.AI20242 cited

GUI Testing Arena: A Unified Benchmark for Advancing Autonomous GUI Testing Agent

Kangjia Zhao, Jiahui Song, Leigang Sha +5

Nowadays, research on GUI agents is a hot topic in the AI community. However, current research focuses on GUI task automation, limiting the scope of applications in various GUI sce…

cs.CV2024

ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Haozhan Shen, Kangjia Zhao, Tiancheng Zhao +4

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in vision-language understanding. Recently, with the integration of test-time scaling techniques,…