14 papers
Ordered Action Tokens for Visuomotor Policy Learning
Chaoqi Liu, Yue Zhao, Haonan Chen +4
Action tokenization maps continuous robot action chunks to discrete tokens and has become an important interface for modern visuomotor policies. Existing approaches either rely on…
Masked Visual Actions for Unified World Modeling
Hadi Alzayer, Wenlong Huang, Haonan Chen +8
Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challe…
DriftWorld: Fast World Modeling through Drifting
Susie Lu, Haonan Chen, Weirui Ye +1
The paper introduces DriftWorld, an action‑conditioned world model that uses a drifting generative approach to produce future frames in a single forward pass, enabling fast (30+ fp…
MemoBench: Benchmarking World Modeling in Dynamically Changing Environments
Haoyu Chen, Kaichen Zhou, Hang Hua +11
The paper introduces MemoBench, a benchmark that tests video generation models' ability to remember and correctly update objects that disappear and later reappear in dynamically ch…
B-spline Policy: Accelerating Manipulation Policies via B-spline Action Representations
Xiaoshen Han, Haoyu Xiong, Haonan Chen +4
In this work, we present B-spline Policy (BSP), an action representation designed for accelerating robot manipulation policies. Rather than predicting discrete-time action chunks,…
ArchEval: Measuring AI Agents as Computer Architects
Chenyu Wang, Zishen Wan, Jeffrey Ma +8
Computer architecture has long used benchmarks to make progress measurable. LLM agents create a different measurement problem: success is not merely writing code or tuning paramete…