5 papers
GeoLanG: Geometry-Aware Language-Guided Grasping with Unified RGB-D Multimodal Learning
Rui Tang, Guankun Wang, Long Bai +6
Language-guided grasping has emerged as a promising paradigm for enabling robots to identify and manipulate target objects through natural language instructions, yet it remains hig…
Unlocking Mixed Reality for Medical Education: A See-Through Perspective on Head Anatomy
Yuqing Wei, Yupeng Wang, Jiayi Zhao +3
Extended reality (XR), encompassing Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR), is emerging as a transformative platform for medical education. Traditiona…
Contact-Aided Navigation of Flexible Robotic Endoscope Using Deep Reinforcement Learning in Dynamic Stomach
Chi Kit Ng, Huxin Gao, Tian-Ao Ren +2
Navigating a flexible robotic endoscope (FRE) through the gastrointestinal tract is critical for surgical diagnosis and treatment. However, navigation in the dynamic stomach is par…
CapsDT: Diffusion-Transformer for Capsule Robot Manipulation
Xiting He, Mingwu Su, Xinqi Jiang +3
Vision-Language-Action (VLA) models have emerged as a prominent research area, showcasing significant potential across a variety of applications. However, their performance in endo…
Multimodal Graph Representation Learning for Robust Surgical Workflow Recognition with Adversarial Feature Disentanglement
Long Bai, Boyi Ma, Ruohan Wang +8
Surgical workflow recognition is vital for automating tasks, supporting decision-making, and training novice surgeons, ultimately improving patient safety and standardizing procedu…