collaborators

19 papers

cs.RO2026

StructRL: Structured Action-Space Exploration for Flow-Based VLAs

Jiarui Yang, Bin Zhu, Jingjing Chen +4

Flow-based Vision-Language-Action (VLA) models are now widely used for continuous robotic manipulation, and online reinforcement learning (RL) is emerging as a key technique for ad…

cs.AI2026

ReGraph: Learning to Generate Recipe Graphs from Food Images

Guoshan Liu, Bin Zhu, Pengkun Jiao +3

Recent Large Multimodal Models (LMMs) have achieved impressive performance in recipe generation from food images.However, cooking is a structured transformation process in which in…

cs.CV2026

Disentangling Semantic Attention from Structural Bias in the Attention Manifold

Pengkun Jiao, Bin Zhu, Jingjing Chen +1

The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws. Specifically, MLLMs consistently exhibit disprop…

cs.RO2026

Restoring Linguistic Grounding in VLA Models via Train-Free Attention Recalibration

Ninghao Zhang, Bin Zhu, Shijie Zhou +1

Vision-Language-Action (VLA) models enable robots to perform manipulation tasks directly from natural language instructions and are increasingly viewed as a foundation for generali…

cs.CV2026

Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models

Ziyao Tang, Pengkun Jiao, Bin Zhu +3

Video Large Language Models (Vid-LLMs) have demonstrated remarkable performance in video understanding tasks, yet their robustness under conversational interaction remains largely…

cs.CV2026

SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning

Yian Li, Yang Jiao, Bin Zhu +4

Spatial intelligence, which refers to the ability to reason about geometric and physical structure from visual observations, remains a core challenge for multimodal large language…