9 papers
ARMFlow: AutoRegressive MeanFlow for Online 3D Human Reaction Generation
Zichen Geng, Zeeshan Hayder, Wei Liu +2
3D human reaction generation faces three main challenges:(1) high motion fidelity, (2) real-time inference, and (3) autoregressive adaptability for online scenarios. Existing metho…
Surg-R1: A Hierarchical Reasoning Foundation Model for Scalable and Interpretable Surgical Decision Support with Multi-Center Clinical Validation
Jian Jiang, Chenxi Lin, Yiming Gu +24
Surgical scene understanding demands not only accurate predictions but also interpretable reasoning that surgeons can verify against clinical expertise. However, existing surgical…
Diffusion Stabilizer Policy for Automated Surgical Robot Manipulations
Chonlam Ho, Jianshu Hu, Lei Song +3
Intelligent surgical robots have the potential to revolutionize clinical practice by enabling more precise and automated surgical procedures. However, the automation of such robot…
4D Monocular Surgical Reconstruction under Arbitrary Camera Motions
Jiwei Shan, Zeyu Cai, Cheng-Tai Hsieh +5
Reconstructing deformable surgical scenes from endoscopic videos is challenging and clinically important. Recent state-of-the-art methods based on implicit neural representations o…
NRGS-SLAM: Monocular Non-Rigid SLAM for Endoscopy via Deformation-Aware 3D Gaussian Splatting
Jiwei Shan, Zeyu Cai, Yirui Li +5
Visual simultaneous localization and mapping (V-SLAM) is a fundamental capability for autonomous perception and navigation. However, endoscopic scenes violate the rigidity assumpti…
FlowVLA: Visual Chain of Thought-based Motion Reasoning for Vision-Language-Action Models
Zhide Zhong, Haodong Yan, Junfeng Li +8
Many Vision-Language-Action (VLA) models are built upon an internal world model trained via next-frame prediction ``''. However, this paradigm attempts to…