9 papers
Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models
Hongyu Li, Wanjia Fu, Xiaoyan Cong +11
Predicting object dynamics (i.e., world modeling) is a fundamental challenge for robotic manipulation, and modeling deformable objects presents a particularly difficult case due to…
EgoTL: Egocentric Think-Aloud Chains for Long-Horizon Tasks
Lulin Liu, Dayou Li, Yiqing Liang +8
Large foundation models have made significant advances in embodied intelligence, enabling synthesis and reasoning over egocentric input for household tasks. However, VLM-based auto…
LASER: Layer-wise Scale Alignment for Training-Free Streaming 4D Reconstruction
Tianye Ding, Yiming Xie, Yiqing Liang +3
Recent feed-forward reconstruction models like VGGT and achieve impressive reconstruction quality but cannot process streaming videos due to quadratic memory complexity, lim…
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not…
E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models
Wenyan Cong, Yiqing Liang, Yancheng Zhang +7
Spatial intelligence, encompassing 3D reconstruction, perception, and reasoning, is fundamental to applications such as robotics, aerial imaging, and extended reality. A key enable…
Monocular Dynamic Gaussian Splatting: Fast, Brittle, and Scene Complexity Rules
Yiqing Liang, Mikhail Okunev, Mikaela Angelina Uy +4
Gaussian splatting methods are emerging as a popular approach for converting multi-view image data into scene representations that allow view synthesis. In particular, there is int…