9 papers
REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak
Jiachen Ma, Jiawen Zhang, Xiangtian Li +3
While Large Language Models (LLMs) demonstrate remarkable capabilities, they remain susceptible to sophisticated, multi-step jailbreak attacks that circumvent conventional surface-…
Agent Skills Should Go Beyond Text: The Case for Visual Skills
Binxiao Xu, Ruichuan An, Bocheng Zou +1
Reusable skills are a key mechanism for extending agent capabilities, allowing agents to accumulate experience and solve increasingly complex tasks. Yet most existing skill-learnin…
ChronoAgentic: A Code-based Multi-Agent World Simulator for Physically Grounded Simulation Construction
Hongyu Wang, Jingquan Wang, Bocheng Zou +3
Video-based world models generate visually plausible rollouts, but since they infer dynamics in latent states, they enforce no explicit physical constraints: contacts drift, shapes…
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
Ziyun Zeng, Hang Hua, Bocheng Zou +3
Recent GUI agents have made substantial progress in visual grounding and action prediction, yet they remain brittle in long-horizon tasks that require maintaining task state across…
Chrono-Gymnasium: An Open-Source, Gymnasium-Compatible Distributed Simulation Framework
Bocheng Zou, Harry Zhang, Khailanii Slaton +5
High-fidelity physics simulation is essential for closing the sim-to-real gap in robotics and complex mechanical systems. However, the computational overhead of high-fidelity engin…
MuRF: Unlocking the Multi-Scale Potential of Vision Foundation Models
Bocheng Zou, Mu Cai, Mark Stanley +2
Vision Foundation Models (VFMs) have become the cornerstone of modern computer vision, offering robust representations across a wide array of tasks. While recent advances allow the…