activity
20232026
most citedVMamba: Visual State Space Model

382 citations · 389 across the 13 of their papers we have counts for

collaborators

13 papers

cs.RO2026

Temporal Forcing: 4D Representation Alignment for Vision-Language-Action Models

Xingyu Ding, Yuzhong Zhao, Chunhai Zhao +3

Recent vision-language-action (VLA) methods improve manipulation performance by aligning their representations with 3D scene geometry. However, these methods often struggle with lo…

cs.RO2026

Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models

Xingyu Ding, Yuzhong Zhao, Yang Wu +4

Recent Vision-Language-Action (VLA) methods improve generalization by aligning their representations with 3D scene geometry. However, these methods are fundamentally instruction-ag…

cs.CL2026

Balancing Understanding and Generation in Discrete Diffusion Models

Yue Liu, Yuzhong Zhao, Zheyong Xie +5

In discrete generative modeling, two dominant paradigms demonstrate divergent capabilities: Masked Diffusion Language Models (MDLM) excel at semantic understanding and zero-shot ge…

cs.CV2025

Thinking with Images via Self-Calling Agent

Wenxi Yang, Yuzhong Zhao, Fang Wan +1

Thinking-with-images paradigms have showcased remarkable visual reasoning capability by integrating visual information as dynamic elements into the Chain-of-Thought (CoT). However,…

cs.CV2025

DocReward: A Document Reward Model for Structuring and Stylizing

Junpeng Liu, Yuzhong Zhao, Bowen Cao +17

Recent agentic workflows automate professional document generation but focus narrowly on textual quality, overlooking structural and stylistic professionalism, which is equally cri…

cs.CL2025

Geometric-Mean Policy Optimization

Yuzhong Zhao, Yue Liu, Junpeng Liu +9

Group Relative Policy Optimization (GRPO) has significantly enhanced the reasoning capability of large language models by optimizing the arithmetic mean of token-level rewards. Unf…