3 papers
cs.RO2026
CometVLA: Co-Training on an Embodied Data Pyramid towards Physical Understanding
Hanwen Wan, Dafeng Chi, Linbo Zhai +6
Vision-language-action (VLA) models remain brittle in manipulation tasks that require physical commonsense. Current physical VQA data is typically disembodied and misaligned with r…
cs.CV2026
ResemBrick: Brick Reconstruction from Photographs with Perceptual Fidelity and Buildability
Xilun Chen, Hanwen Wan, Yusong Zhao +3
Producing a hand-buildable, colored brick model of a 3D object from a few casual photographs is a clean testbed for a broader challenge: generating 3D content that meets hard physi…
cs.RO2026
JoyAI-RA 0.1: A Foundation Model for Robotic Autonomy
Tianle Zhang, Zhihao Yuan, Dafeng Chi +59
Robotic autonomy in open-world environments is fundamentally limited by insufficient data diversity and poor cross-embodiment generalization. Existing robotic datasets are often li…