9 papers
Enabling Extensible Embodied Capabilities with Tools
Xueyang Zhou, Zijia Wang, Qianjiang Li +7
Most existing embodied intelligence methods formulate perception, reasoning, planning, and control within a unified parameterized policy. Yet these capabilities are inherently hier…
LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
Xueyang Zhou, Yangming Xu, Guiyao Tie +5
LIBERO has emerged as a widely adopted benchmark for evaluating Vision-Language-Action (VLA) models; however, its current training and evaluation settings are problematic, often le…
AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery
Guiyao Tie, Jiawen Shi, Dingjie Song +20
Scientific research is being reshaped by AI systems that move beyond isolated assistance toward longer-horizon workflows spanning literature grounding, hypothesis generation, exper…
EmbodiedClaw: Conversational Workflow Execution for Embodied AI Development
Xueyang Zhou, Yihan Sun, Xijie Gong +4
Embodied AI research is increasingly moving beyond single-task, single-environment policy learning toward multi-task, multi-scene, and multi-model settings. This shift substantiall…
SafeAgent: Safeguarding LLM Agents via an Automated Risk Simulator
Xueyang Zhou, Weidong Wang, Lin Lu +7
Large Language Model (LLM)-based agents are increasingly deployed in real-world applications such as "digital assistants, autonomous customer service, and decision-support systems"…
MMLU-Reason: Benchmarking Multi-Task Multi-modal Language Understanding and Reasoning
Guiyao Tie, Xueyang Zhou, Tianhe Gu +7
Recent advances in Multi-Modal Large Language Models (MLLMs) have enabled unified processing of language, vision, and structured inputs, opening the door to complex tasks such as l…