2 papers
cs.RO2025
CapsDT: Diffusion-Transformer for Capsule Robot Manipulation
Xiting He, Mingwu Su, Xinqi Jiang +3
Vision-Language-Action (VLA) models have emerged as a prominent research area, showcasing significant potential across a variety of applications. However, their performance in endo…
cs.CV2025
EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery
Guankun Wang, Long Bai, Junyi Wang +13
Recently, Multimodal Large Language Models (MLLMs) have demonstrated their immense potential in computer-aided diagnosis and decision-making. In the context of robotic-assisted sur…