6 papers
Track4Action: Distilling World-Centric 3D Tracker into Vision-Language-Action Policies
Chenyi Wang, Xinkai Wang, Bokai Lin +4
Action labels tell a vision-language-action (VLA) policy which robot commands to imitate, but not how those commands change the 3D world. The aligned demonstration clip contains th…
LaMP: Learning Vision-Language-Action Policy with 3D Scene Flow as Latent Motion Prior
Xinkai Wang, Chenyi Wang, Yifu Xu +7
We introduce \textbf{LaMP}, a dual-expert Vision-Language-Action framework that embeds dense 3D scene flow as a latent motion prior for robotic manipulation.Existing VLA models reg…
Impact of Target and Tool Visualization on Depth Perception and Usability in Optical See-Through AR
Yue Yang, Xue Xie, Xinkai Wang +8
Optical see-through augmented reality (OST-AR) systems like Microsoft HoloLens 2 hold promise for arm's distance guidance (e.g., surgery), but depth perception of the hologram and…
Interaction as Intelligence: Deep Research With Human-AI Partnership
Lyumanshan Ye, Xiaojie Cai, Xinkai Wang +23
This paper introduces "Interaction as Intelligence" research series, presenting a reconceptualization of human-AI relationships in deep research tasks. Traditional approaches treat…
MRUCT: Mixed Reality Assistance for Acupuncture Guided by Ultrasonic Computed Tomography
Xinkai Wang, Yue Yang, Kehong Zhou +4
Chinese acupuncture practitioners primarily depend on muscle memory and tactile feedback to insert needles and accurately target acupuncture points, as the current workflow lacks i…
DipMe: Haptic Recognition of Granular Media for Tangible Interactive Applications
Xinkai Wang, Shuo Zhang, Ziyi Zhao +2
While tangible user interface has shown its power in naturally interacting with rigid or soft objects, users cannot conveniently use different types of granular materials as the in…