9 papers
Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models
Xingyu Ding, Yuzhong Zhao, Yang Wu +4
Recent Vision-Language-Action (VLA) methods improve generalization by aligning their representations with 3D scene geometry. However, these methods are fundamentally instruction-ag…
DocReward: A Document Reward Model for Structuring and Stylizing
Junpeng Liu, Yuzhong Zhao, Bowen Cao +17
Recent agentic workflows automate professional document generation but focus narrowly on textual quality, overlooking structural and stylistic professionalism, which is equally cri…
Balancing Understanding and Generation in Discrete Diffusion Models
Yue Liu, Yuzhong Zhao, Zheyong Xie +5
In discrete generative modeling, two dominant paradigms demonstrate divergent capabilities: Masked Diffusion Language Models (MDLM) excel at semantic understanding and zero-shot ge…
Thinking with Images via Self-Calling Agent
Wenxi Yang, Yuzhong Zhao, Fang Wan +1
Thinking-with-images paradigms have showcased remarkable visual reasoning capability by integrating visual information as dynamic elements into the Chain-of-Thought (CoT). However,…
Geometric-Mean Policy Optimization
Yuzhong Zhao, Yue Liu, Junpeng Liu +9
Group Relative Policy Optimization (GRPO) has significantly enhanced the reasoning capability of large language models by optimizing the arithmetic mean of token-level rewards. Unf…
Model as a Game: On Numerical and Spatial Consistency for Generative Games
Jingye Chen, Yuzhong Zhao, Yupan Huang +5
Recent advances in generative models have significantly impacted game generation. However, despite producing high-quality graphics and adequately receiving player input, existing m…