collaborators

5 papers

cs.CV2026

Advancing MLLM-based UAV Image Understanding and Reasoning: A Benchmark and a Training-Free Multi-Agent System

Haoyu Zhang, Shuoxun Zhang, Peng Ye +5

Multimodal Large Language Model (MLLM)-based UAV aerial image understanding and reasoning is essential for aerial intelligence yet poses distinct challenges arising from extreme sc…

cs.RO2026

Trust Your Instincts: Confidence-Driven Test-Time RL for Vision-Language-Action Models

Siyao Chen, Jiakang Yuan, Jiaxin Wang +1

Reinforcement learning (RL) has become indispensable for pushing Vision-Language-Action Models (VLAs) beyond static imitation learning. However, existing RL methods typically requi…

cs.CV2026

FreshMem: Brain-Inspired Frequency-Space Hybrid Memory for Streaming Video Understanding

Kangcong Li, Peng Ye, Lin Zhang +3

Transitioning Multimodal Large Language Models (MLLMs) from offline to online streaming video understanding is essential for continuous perception. However, existing methods lack f…

cs.CV2025

Sequential Token Merging: Revisiting Hidden States

Yan Wen, Peng Ye, Lin Zhang +4

Vision Mambas (ViMs) achieve remarkable success with sub-quadratic complexity, but their efficiency remains constrained by quadratic token scaling with image resolution. While exis…

cs.CV2025

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning

Lin Zhang, Xianfang Zeng, Kangcong Li +2

We propose SC-Captioner, a reinforcement learning framework that enables the self-correcting capability of image caption models. Our crucial technique lies in the design of the rew…