collaborators

10 papers

cs.CV2026

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling

Qu Tang, Benhui Zhuang, Bo Yuan +3

Vision-language-action (VLA) models are a widely adopted paradigm for embodied policies. They excel at efficient closed-loop control but do not explicitly model how physical scenes…

cs.AI2026

SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction

Xue Yu, Bo Yuan, Pengshuai Yang +3

Mobile graphical user interface (GUI) agents have demonstrated remarkable capabilities in automating complex tasks, yet they introduce critical safety risks where a single erroneou…

cs.LG2026

Enhancing Reasoning for Diffusion LLMs via Distribution Matching Policy Optimization

Yuchen Zhu, Wei Guo, Jaemoo Choi +4

Diffusion large language models (dLLMs) are promising alternatives to autoregressive large language models (AR-LLMs), as they potentially allow higher inference throughput. Reinfor…

cs.CV2026

WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling

Zelin Zhao, Min Shi, Bo Yuan +5

World models aim to capture environment dynamics in ways that support perception, reasoning, and action, and have recently become a central direction in Vision-Language-Action-Worl…

cs.MA2026

PersonalPlan: Planning Multi-Agent Systems for Personalized Programming Learning

Zhiyuan Wen, Jiannong Cao, Peng Gao +4

Effective programming education requires personalized instruction adapted to diverse learner backgrounds. However, while LLM-based multi-agent systems (MAS) excel at complex planni…

cs.LG2026

Clarify Before You Draw: Proactive Agents for Robust Text-to-CAD Generation

Bo Yuan, Zelin Zhao, Petr Molodyk +2

Large language models have recently enabled text-to-CAD systems that synthesize parametric CAD programs (e.g., CadQuery) from natural-language prompts. In practice, however, geomet…