activity
20242026
most citedMarineEval: Assessing the Marine Intelligence of Vision-Language Models

1 citations · 1 across the 15 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

RS-HyRe-R1: A Hybrid Reward Mechanism to Overcome Perceptual Inertia for Remote Sensing Images Understanding

Gaozhi Zhou, Hu He, Peng Shen +8

Reinforcement learning (RL) post-training substantially improves remote sensing vision-language models (RS-VLMs). However, when handling complex remote sensing imagery (RSI) requir…

cs.CV2026

FarmMind: Reasoning-Query-Driven Dynamic Segmentation for Farmland Remote Sensing Images

Haiyang Wu, Weiliang Mu, Jipeng Zhang +4

Existing methods for farmland remote sensing image (FRSI) segmentation generally follow a static segmentation paradigm, where analysis relies solely on the limited information cont…

cs.CV2026

MMedExpert-R1: Strengthening Multimodal Medical Reasoning via Domain-Specific Adaptation and Clinical Guideline Reinforcement

Meidan Ding, Jipeng Zhang, Wenxuan Wang +4

Medical Vision-Language Models (MedVLMs) excel at perception tasks but struggle with complex clinical reasoning required in real-world scenarios. While reinforcement learning (RL)…

cs.CV20251 cited

MarineEval: Assessing the Marine Intelligence of Vision-Language Models

YuK-Kwan Wong, Tuan-An To, Jipeng Zhang +2

We have witnessed promising progress led by large language models (LLMs) and further vision language models (VLMs) in handling various queries as a general-purpose assistant. VLMs,…

cs.CV2025

Skin-R1: Clinical Knowledge-Guided Dermatological Diagnosis Using Vision-Language Models

Zehao Liu, Weijieying Ren, Jipeng Zhang +4

Vision--language models (VLMs) have recently shown promise for assisting clinical reasoning in dermatological diagnosis. However, their trustworthiness and clinical utility remain…

cs.CV2025

Pointing to a Llama and Call it a Camel: On the Sycophancy of Multimodal Large Language Models

Renjie Pi, Kehao Miao, Li Peihang +4

Multimodal large language models (MLLMs) have demonstrated extraordinary capabilities in conducting conversations based on image inputs. However, we observe that MLLMs exhibit a pr…