collaborators

5 papers

cs.CV2026

EgoAfford: Task-Oriented Affordance Grounding via Egocentric Referring Segmentation

Xinyuan Guan, Feifan Chen, Xinyu Zhan +3

Part-level affordance grounding has advanced the localization of functional object regions associated with elemental actions. Extending this capability to complex tasks calls for c…

cs.RO2026

Track4Action: Distilling World-Centric 3D Tracker into Vision-Language-Action Policies

Chenyi Wang, Xinkai Wang, Bokai Lin +4

Action labels tell a vision-language-action (VLA) policy which robot commands to imitate, but not how those commands change the 3D world. The aligned demonstration clip contains th…

cs.RO2026

ChronoFlow-Policy: Unifying Past-Current-Future Interaction Flow in Visuomotor Policy Learning

Bokai Lin, Yifu Xu, Xinyu Zhan +6

Visual signals play a crucial role in policy learning by enabling models to capture object motion and interaction dynamics. Just as humans reason about actions using both past expe…

cs.CV2026

LaMP: Learning Vision-Language-Action Policy with 3D Scene Flow as Latent Motion Prior

Xinkai Wang, Chenyi Wang, Yifu Xu +7

We introduce \textbf{LaMP}, a dual-expert Vision-Language-Action framework that embeds dense 3D scene flow as a latent motion prior for robotic manipulation.Existing VLA models reg…

cs.CV2025

ArtGen: Conditional Generative Modeling of Articulated Objects in Arbitrary Part-Level States

Haowen Wang, Xiaoping Yuan, Fugang Zhang +4

Generating articulated assets is crucial for robotics, digital twins, and embodied intelligence. Existing generative models often rely on single-view inputs representing closed sta…