2 papers
cs.AI2026
ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents
Binjie Zhang, Mike Zheng Shou
Tool-augmented vision-language models (VLMs) can solve multimodal, multi-step tasks by calling external tools, yet they remain fragile in practice. Existing works have two common g…
cs.CV2025
Ego-centric Predictive Model Conditioned on Hand Trajectories
Binjie Zhang, Mike Zheng Shou
In egocentric scenarios, anticipating both the next action and its visual outcome is essential for understanding human-object interactions and for enabling robotic planning. Howeve…