collaborators

14 papers

cs.RO2026

Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents

Guanxiong Chen, Qianjun Xia, Jiawei Peng +21

Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a streamlined real2sim process must recover…

cs.CV2026

FreeStyle: Free Control of Style-Content Dual-Reference Generation from Community LoRA Mining

Jinghong Lan, Wei Cheng, Yunuo Chen +10

Style-content dual-reference generation aims to synthesize an image that preserves the structure and semantics of a content reference while adopting the style of a separate style r…

cs.RO2026

MemoryVAM: Integrating Memory into Video Action Model for Robot Manipulation

Yuxin Jiang, Chang Yu, Yunuo Chen +4

Video-world-model policies learn action-relevant representations by predicting future observations. However, they condition on only a short observation window, which renders long-h…

cs.RO2026

Sparse2Act: Learning Action-Aligned Sparse 3D Representations for Cross-Domain Robot Manipulation

Yu Guo, Chang Yu, Siyu Ma +4

Explicit 3D representations are attractive for manipulation because they expose object shape, workspace geometry, and robot-object relations in metric coordinates. However, sparse…

cs.RO2026

TacCoRL: Integrating Tactile Feedback into VLA via Simulation

Siyu Ma, Yuqi Liang, Chang Yu +5

Vision-language-action (VLA) models provide strong visual, language, and action priors for robot manipulation, but visual observations alone often miss the local contact state requ…

cs.GR2026

Learn2Fold: Structured Origami Generation with World Model Planning

Yanjia Huang, Yunuo Chen, Ying Jiang +4

The ability to transform a flat sheet into a complex three-dimensional structure is a fundamental test of physical intelligence. Unlike cloth manipulation, origami is governed by s…