3 papers
cs.RO2026
Explicit Kinematic Guidance from Analytic Concepts for Vision-Language-Action Models
Mingyang Sun, Jiude Wei, Xiujian Liang +4
Current Vision-Language-Action (VLA) models rely mainly on 2D inputs, neglecting the rich object structural information and commonsense knowledge inherent in the 3D physical world.…
cs.RO2026
Analytic Concept-Centric Memory for Agentic Embodied Manipulation
Mingyang Sun, Xiujian Liang, Jiude Wei +4
Long-horizon embodied manipulation requires agents to remember persistent objects, track changing scene states, and reuse prior interaction knowledge. However, existing agent memor…
cs.RO2025
RealDiff: Bridging Real-World Gap in Robot Manipulation via Depth Diffusion
Xiujian Liang, Jiacheng Liu, Mingyang Sun +3
Robot manipulation in the real world is fundamentally constrained by the visual sim2real gap, where depth observations collected in simulation fail to reflect the complex noise pat…