collaborators

32 papers

cs.CV2026

ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning

Ting Huang, Zhenyu Zhang, Wenyuan Huang +2

Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where models must infer spatial relations across long horizons under chan…

cs.CV2026

WebCryptoAgent: Agentic Crypto Trading with Web Informatics

Ali Kurban, Wei Luo, Liangyu Zuo +5

Cryptocurrency trading increasingly depends on timely integration of heterogeneous web information and market microstructure signals to support short-horizon decision making under…

cs.CV2026

GeneralVLA-2: Geometry-Aware Reconstruction and Governed Memory for Robot Planning

Haoyu Wang, Guoqing Ma, Zeyu Zhang +3

Generalist vision-language-action systems need object-centric 3D evidence and reusable manipulation experience to plan reliable robot trajectories. GeneralVLA provides a hierarchic…

cs.RO2026

Learning New Tasks via Reusable Skills: Skill-Compositional Experts for Embodied Continual Learning

Shuaike Zhang, Shaokun Wang, Haoyu Tang +2

Embodied Continual Learning (ECL) aims to enable robots to continually acquire new manipulation tasks while retaining previously learned behaviors under closed-loop control. Compar…

cs.CV2026

MotionVLA: Vision-Language-Action Model for Humanoid Motion

Nonghai Zhang, Siyu Zhai, Yanjun Li +5

Generating realistic humanoid motion from scene images and text involves both low-frequency pose semantics and high-frequency physical dynamics. However, many existing methods toke…

cs.RO2026

DragMesh-2: Physically Plausible Dexterous Hand-Object Interaction with Articulated Objects

Tianshan Zhang, Yijia Duan, Yanjun Li +2

Dexterous interaction with articulated objects is important for household, assistive, and humanoid manipulation, where multi-finger hands can provide compliant contact patterns bey…