9 papers
MemoryVAM: Integrating Memory into Video Action Model for Robot Manipulation
Yuxin Jiang, Chang Yu, Yunuo Chen +4
Video-world-model policies learn action-relevant representations by predicting future observations. However, they condition on only a short observation window, which renders long-h…
Sparse2Act: Learning Action-Aligned Sparse 3D Representations for Cross-Domain Robot Manipulation
Yu Guo, Chang Yu, Siyu Ma +4
Explicit 3D representations are attractive for manipulation because they expose object shape, workspace geometry, and robot-object relations in metric coordinates. However, sparse…
Fishbone: From One 3D Asset to a Million Controllable Edits
Yumeng He, Xiaoying Wang, Peihao Li +7
Large-scale controllable 3D assets are critical for computer graphics, embodied AI, robotics, and interactive content creation, yet creating diverse 3D assets remains challenging d…
Gradient Descent with Projection Finds Over-Parameterized Neural Networks for Learning Low-Degree Polynomials with Nearly Minimax Optimal Rate
Yingzhen Yang, Ping Li
We study the problem of learning a low-degree spherical polynomial of degree defined on the unit sphere in $\RR^d$ by training an over-parameterized two-layer n…
VoroLight: Learning Voronoi Surface Meshes via Sphere Intersection
Jiayin Lu, Ying Jiang, Yumeng He +2
Voronoi diagrams naturally produce convex, watertight, and topologically consistent cells, making them an appealing representation for 3D shape reconstruction. However, standard di…
SeeClear: Reliable Transparent Object Depth Estimation via Generative Opacification
Xiaoying Wang, Yumeng He, Jingkai Shi +4
Monocular depth estimation remains challenging for transparent objects, where refraction and transmission are difficult to model and break the appearance assumptions used by depth…