collaborators

9 papers

cs.CV2026

MedSAM-Agent: Empowering Interactive Medical Image Segmentation with Multi-turn Agentic Reinforcement Learning

Shengyuan Liu, Liuxin Bao, Qi Yang +6

Medical image segmentation is evolving from task-specific models toward generalizable frameworks. Recent research leverages Multi-modal Large Language Models (MLLMs) as autonomous…

cs.LG2026

Dual-Phase LLM Reasoning: Self-Evolved Mathematical Frameworks

ShaoZhen Liu, Xinting Huang, Houwen Peng +4

In recent years, large language models (LLMs) have demonstrated significant potential in complex reasoning tasks like mathematical problem-solving. However, existing research predo…

cs.CV2025

HunyuanOCR Technical Report

Hunyuan Vision Team, Pengyuan Lyu, Xingyu Wan +29

This paper presents HunyuanOCR, a commercial-grade, open-source, and lightweight (1B parameters) Vision-Language Model (VLM) dedicated to OCR tasks. The architecture comprises a Na…

cs.CV2025

R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning

Qi Yang, Bolin Ni, Shiming Xiang +3

Multimodal Large Language Models (MLLMs) equipped with step-by-step thinking capabilities have demonstrated remarkable performance on complex reasoning problems. However, this thin…

cs.CV2025

RDPO: Real Data Preference Optimization for Physics Consistency Video Generation

Wenxu Qian, Chaoyue Wang, Hou Peng +3

Video generation techniques have achieved remarkable advancements in visual quality, yet faithfully reproducing real-world physics remains elusive. Preference-based model post-trai…

cs.CV2025

RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs

Meng-Hao Guo, Xuanyu Chu, Qianrui Yang +12

The rapid advancement of native multi-modal models and omni-models, exemplified by GPT-4o, Gemini, and o3, with their capability to process and generate content across modalities s…