activity
20242026
collaborators

9 papers

cs.LG2026

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak

Jiachen Ma, Jiawen Zhang, Xiangtian Li +3

While Large Language Models (LLMs) demonstrate remarkable capabilities, they remain susceptible to sophisticated, multi-step jailbreak attacks that circumvent conventional surface-…

cs.CV2026

Agent Skills Should Go Beyond Text: The Case for Visual Skills

Binxiao Xu, Ruichuan An, Bocheng Zou +1

Reusable skills are a key mechanism for extending agent capabilities, allowing agents to accumulate experience and solve increasingly complex tasks. Yet most existing skill-learnin…

cs.AI2026

ChronoAgentic: A Code-based Multi-Agent World Simulator for Physically Grounded Simulation Construction

Hongyu Wang, Jingquan Wang, Bocheng Zou +3

Video-based world models generate visually plausible rollouts, but since they infer dynamics in latent states, they enforce no explicit physical constraints: contacts drift, shapes…

cs.CV2026

MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents

Ziyun Zeng, Hang Hua, Bocheng Zou +3

Recent GUI agents have made substantial progress in visual grounding and action prediction, yet they remain brittle in long-horizon tasks that require maintaining task state across…

cs.RO2026

Chrono-Gymnasium: An Open-Source, Gymnasium-Compatible Distributed Simulation Framework

Bocheng Zou, Harry Zhang, Khailanii Slaton +5

High-fidelity physics simulation is essential for closing the sim-to-real gap in robotics and complex mechanical systems. However, the computational overhead of high-fidelity engin…

cs.CV2026

MuRF: Unlocking the Multi-Scale Potential of Vision Foundation Models

Bocheng Zou, Mu Cai, Mark Stanley +2

Vision Foundation Models (VFMs) have become the cornerstone of modern computer vision, offering robust representations across a wide array of tasks. While recent advances allow the…