Publications (116)
SegSTRONG-C: Segmenting Surgical Tools Robustly On Non-adversarial Generated Corruptions -- An EndoVis'24 Challenge
Hao Ding, Yuqian Zhang, Tuxun Lu +39
Surgical data science has seen rapid advancement with the excellent performance of end-to-end deep neural networks (DNNs). Despite their successes, DNNs have been proven susceptibl…
What is Event Knowledge Graph: A Survey
Saiping Guan, Xueqi Cheng, Long Bai +5
Besides entity-centric knowledge, usually organized as Knowledge Graph (KG), events are also an essential kind of knowledge in the world, which trigger the spring up of event-centr…
Extremes of Gaussian chaos processes with Trend
Long Bai
Let be a Gaussian vector process and let be a continuous homogeneous functi…
CoPESD: A Multi-Level Surgical Motion Dataset for Training Large Vision-Language Models to Co-Pilot Endoscopic Submucosal Dissection
Guankun Wang, Han Xiao, Huxin Gao +6
submucosal dissection (ESD) enables rapid resection of large lesions, minimizing recurrence rates and improving long-term overall survival. Despite these advantages, ESD is technic…
Nested Event Extraction upon Pivot Element Recogniton
Weicheng Ren, Zixuan Li, Xiaolong Jin +6
Nested Event Extraction (NEE) aims to extract complex event structures where an event contains other events as its arguments recursively. Nested events involve a kind of Pivot Elem…
Towards Robust Universal Information Extraction: Benchmark, Evaluation, and Solution
Jizhao Zhu, Akang Shi, Zixuan Li +4
In this paper, we aim to enhance the robustness of Universal Information Extraction (UIE) by introducing a new benchmark dataset, a comprehensive evaluation, and a feasible solutio…
The Influence of Age and Gender Information on the Diagnosis of Diabetic Retinopathy: Based on Neural Networks
Long Bai, Sihang Chen, Mingyang Gao +3
This paper proposes the importance of age and gender information in the diagnosis of diabetic retinopathy. We utilized Deep Residual Neural Networks (ResNet) and Densely Connected…
Extremes of -norm of Vector-valued Gaussian processes with Trend
Long Bai
Let be a Gaussian vector process and be a continuous function. The asymptotics of distribution of $\left\|\boldsymbol{X}(t)\right\…
DaX: Learning General Pathology Representations Across Scales
Bokai Zhao, Yiyang Zhang, Long Bai +3
Computational pathology requires visual representations that transfer across diverse clinical endpoints and remain robust to variation in magnification, staining, scanner type, sli…
GeoLanG: Geometry-Aware Language-Guided Grasping with Unified RGB-D Multimodal Learning
Rui Tang, Guankun Wang, Long Bai +6
Language-guided grasping has emerged as a promising paradigm for enabling robots to identify and manipulate target objects through natural language instructions, yet it remains hig…
TMR-VLA:Vision-Language-Action Model for Magnetic Motion Control of Tri-leg Silicone-based Soft Robot
Ruijie Tang, Chi Kit Ng, Kaixuan Wu +5
In-vivo environments, magnetically actuated soft robots offer advantages such as wireless operation and precise control, showing promising potential for painless detection and ther…
Parisian ruin of Brownian motion risk model over an infinite-time horizon
Long Bai
Let be a standard Brownian motion. In this paper, we derive the exact asymptotics of the probability of Parisian ruin on infinite time horizon for the follo…
FinDeepIndicator: Benchmarking Deep Research Agents in End-to-End Financial Indicator Construction
Chaoqun Yang, Fengbin Zhu, Xinyu Lin +5
Financial indicators are essential tools for transforming raw financial data into interpretable measures for various downstream tasks, such as valuation, risk assessment, and econo…
SurgOnAir: Hierarchy-Aware Real-Time Surgical Video Commentary
Jingyi He, Yue Zhou, Long Bai +3
Understanding surgical workflow in real time is fundamental for intelligent surgical embodiment, where AI systems continuously perceive and respond as surgery proceeds. In the oper…
Rich Event Modeling for Script Event Prediction
Long Bai, Saiping Guan, Zixuan Li +3
Script is a kind of structured knowledge extracted from texts, which contains a sequence of events. Based on such knowledge, script event prediction aims to predict the subsequent…
Event Ontology Expansion via LLM-Based Conceptualization
Weicheng Ren, Zixuan Li, Long Bai +3
Event ontology expansion aims to discover emerging event types from data and extend them to appropriate positions in the existing event ontology.. Existing methods typically cluste…
Surg-UniWorld: A Unified Surgical World Model with Multimodal Control Experts
Rulin Zhou, Wanhao Liu, Guoheng Ma +8
Controllable surgical world models can provide a generative foundation for surgical artificial intelligence and simulation by synthesizing realistic instrument--tissue interactions…
Surgical-VQLA++: Adversarial Contrastive Learning for Calibrated Robust Visual Question-Localized Answering in Robotic Surgery
Long Bai, Guankun Wang, Mobarakol Islam +3
Medical visual question answering (VQA) bridges the gap between visual information and clinical decision-making, enabling doctors to extract understanding from clinical images and…
Integrating Deep Event-Level and Script-Level Information for Script Event Prediction
Long Bai, Saiping Guan, Jiafeng Guo +3
Scripts are structured sequences of events together with the participants, which are extracted from the texts.Script event prediction aims to predict the subsequent event given the…
Endo-4DGX: Robust Endoscopic Scene Reconstruction and Illumination Correction with Gaussian Splatting
Yiming Huang, Long Bai, Beilei Cui +7
Accurate reconstruction of soft tissue is crucial for advancing automation in image-guided robotic surgery. The recent 3D Gaussian Splatting (3DGS) techniques and their variants, 4…
Extremes of vector-valued Gaussian processes with Trend
Long Bai, Krzysztof Debicki, Peng Liu
Let be a centered vector-valued Gaussian process with independent components and continuous trajectories, and $h…
SurgPLAN++: Universal Surgical Phase Localization Network for Online and Offline Inference
Zhen Chen, Xingjian Luo, Jinlin Wu +5
Surgical phase recognition is critical for assisting surgeons in understanding surgical videos. Existing studies focused more on online surgical phase recognition, by leveraging pr…
Extremes of Gaussian random fields with non-additive dependence structure
Long Bai, Krzysztof Debicki, Peng Liu
We derive exact asymptotics of for a centered Gaussian field $X(\mathbf{t}),~…
EndoOOD: Uncertainty-aware Out-of-distribution Detection in Capsule Endoscopy Diagnosis
Qiaozhi Tan, Long Bai, Guankun Wang +2
Wireless capsule endoscopy (WCE) is a non-invasive diagnostic procedure that enables visualization of the gastrointestinal (GI) tract. Deep learning-based methods have shown effect…
Extremes of -Locally stationary Gaussian processes with non-constant variances
Long Bai
With motivation from K. DÈ©bicki and P. Kisowski (2007), in this paper we derive the exact tail asymptotics of -locally stationary Gaussian processes with non-constant varia…
PvNeXt: Rethinking Network Design and Temporal Motion for Point Cloud Video Recognition
Jie Wang, Tingfa Xu, Lihe Ding +3
Point cloud video perception has become an essential task for the realm of 3D vision. Current 4D representation learning techniques typically engage in iterative processing coupled…
Semi-supervised Learning for Segmentation of Bleeding Regions in Video Capsule Endoscopy
Hechen Li, Yanan Wu, Long Bai +3
In the realm of modern diagnostic technology, video capsule endoscopy (VCE) is a standout for its high efficacy and non-invasive nature in diagnosing various gastrointestinal (GI)…
Where It Moves, It Matters: Referring Surgical Instrument Segmentation via Motion
Meng Wei, Kun Yuan, Shi Li +7
Enabling intuitive, language-driven interaction with surgical scenes is a critical step toward intelligent operating rooms and autonomous surgical robotic assistance. However, the…
ASI-Seg: Audio-Driven Surgical Instrument Segmentation with Surgeon Intention Understanding
Zhen Chen, Zongming Zhang, Wenwu Guo +5
Surgical instrument segmentation is crucial in surgical scene understanding, thereby facilitating surgical safety. Existing algorithms directly detected all instruments of pre-defi…
On generalized Piterbarg-Berman function
Chengxiu Ling, Hong Zhang, Long Bai
This paper aims to evaluate the Piterbarg-Berman function given by $$\mathcal{P\!B}_α^h(x, E) = \int_\mathbb{R}e^z\mathbb{P} \left\{{\int_E \mathbb{I}\left(\sqrt2B_α(t) - |t|^α-…
Does DINOv3 Set a New Medical Vision Standard? Benchmarking 2D and 3D Classification, Segmentation, and Registration
Che Liu, Yinda Chen, Haoyuan Shi +21
The advent of large-scale vision foundation models, pre-trained on diverse natural images, has marked a paradigm shift in computer vision. However, how the frontier vision foundati…
Drawdown and drawup for fractional Brownian motion with trend
Long Bai, Peng Liu
In this paper, we consider the drawdown and drawup of the fractional Brownian motion with trend, which corresponds to the logarithm of geometric fractional Brownian motion represen…
Endo-4DGS: Endoscopic Monocular Scene Reconstruction with 4D Gaussian Splatting
Yiming Huang, Beilei Cui, Long Bai +4
In the realm of robot-assisted minimally invasive surgery, dynamic scene reconstruction can significantly enhance downstream tasks and improve surgical outcomes. Neural Radiance Fi…
Domain Adaptive Sim-to-Real Segmentation of Oropharyngeal Organs
Guankun Wang, Tian-Ao Ren, Jiewen Lai +2
Video-assisted transoral tracheal intubation (TI) necessitates using an endoscope that helps the physician insert a tracheal tube into the glottis instead of the esophagus. The gro…
Ruin problem for Brownian motion risk model with interest rate and tax payment
Long Bai, Peng Liu
Let be a Brownian motion. Consider the Brownian motion risk model with interest rate collection and tax payment defined by \begin{align}\label{Rudef} \widetilde{…
Can DeepSeek Reason Like a Surgeon? An Empirical Evaluation for Vision-Language Understanding in Robotic-Assisted Surgery
Boyi Ma, Yanguang Zhao, Jie Wang +5
The DeepSeek models have shown exceptional performance in general scene understanding, question-answering (QA), and text generation tasks, owing to their efficient training paradig…
Class-Incremental Few-Shot Event Detection
Kailin Zhao, Xiaolong Jin, Long Bai +2
Event detection is one of the fundamental tasks in information extraction and knowledge graph. However, a realistic event detection system often needs to deal with new event classe…
SAM 2 in Robotic Surgery: An Empirical Evaluation for Robustness and Generalization in Surgical Video Segmentation
Jieming Yu, An Wang, Wenzhen Dong +5
The recent Segment Anything Model (SAM) 2 has demonstrated remarkable foundational competence in semantic segmentation, with its memory mechanism and mask decoder further addressin…
An In-Context Schema Understanding Method for Knowledge Base Question Answering
Yantao Liu, Zixuan Li, Xiaolong Jin +5
The Knowledge Base Question Answering (KBQA) task aims to answer natural language questions based on a given knowledge base. Recently, Large Language Models (LLMs) have shown stron…
Multimodal Graph Representation Learning for Robust Surgical Workflow Recognition with Adversarial Feature Disentanglement
Long Bai, Boyi Ma, Ruohan Wang +8
Surgical workflow recognition is vital for automating tasks, supporting decision-making, and training novice surgeons, ultimately improving patient safety and standardizing procedu…
How can reasoning capability empower the AI copilot robot in endoscopic surgery
Guankun Wang, Long Bai, Hongliang Ren
Reasoning capability has significantly advanced complex logical inference and robotic decision-making in general domains. However, its potential in the Artificial Intelligence (AI)…
V-SfMLearner: Learning Monocular Depth and Ego-motion for Multimodal Wireless Capsule Endoscopy
Long Bai, Beilei Cui, Liangyu Wang +9
Deep learning can predict depth maps and capsule ego-motion from capsule endoscopy videos, aiding in 3D scene reconstruction and lesion localization. However, the collisions of the…
LLCaps: Learning to Illuminate Low-Light Capsule Endoscopy with Curved Wavelet Attention and Reverse Diffusion
Long Bai, Tong Chen, Yanan Wu +3
Wireless capsule endoscopy (WCE) is a painless and non-invasive diagnostic tool for gastrointestinal (GI) diseases. However, due to GI anatomical constraints and hardware manufactu…
A Review of 3D Reconstruction Techniques for Deformable Tissues in Robotic Surgery
Mengya Xu, Ziqi Guo, An Wang +2
As a crucial and intricate task in robotic minimally invasive surgery, reconstructing surgical scenes using stereo or monocular endoscopic video holds immense potential for clinica…
Parisian Ruin of the Brownian Motion Risk Model with Constant Force of Interest
Long Bai, Li Luo
Let be a standard Brownian motion. Define a risk process \label{Rudef} R_u^δ(t)=e^{δt}\left(u+c\int^{t}_{0}e^{-δs}d s-Ï\int_{0}^{t}e^{-δs}d B(s)\right)…
Comparative validation of surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation in endoscopy: Results of the PhaKIR 2024 challenge
Tobias Rueckert, David Rauber, Raphaela Maerkl +58
Reliable recognition and localization of surgical instruments in endoscopic video recordings are foundational for a wide range of applications in computer- and robot-assisted minim…
SAR-RARP50: Segmentation of surgical instrumentation and Action Recognition on Robot-Assisted Radical Prostatectomy Challenge
Dimitrios Psychogyios, Emanuele Colleoni, Beatrice Van Amsterdam +47
Surgical tool segmentation and action recognition are fundamental building blocks in many computer-assisted intervention applications, ranging from surgical skills assessment to de…
EndoGSim: Physics-Aware 4D Dynamic Endoscopic Scene Simulations via MLLM-Guided Gaussian Splatting
Changjing Liu, Yiming Huang, Long Bai +2
In robot-assisted minimally invasive surgery, high-fidelity dynamic endoscopic scene reconstruction and simulation are crucial to enhancing downstream tasks and advancing surgical…
SurgSora: Object-Aware Diffusion Model for Controllable Surgical Video Generation
Tong Chen, Shuya Yang, Junyi Wang +3
Surgical video generation can enhance medical education and research, but existing methods lack fine-grained motion control and realism. We introduce SurgSora, a framework that gen…
Towards Event Extraction with Massive Types: LLM-based Collaborative Annotation and Partitioning Extraction
Wenxuan Liu, Zixuan Li, Long Bai +5
Developing a general-purpose extraction system that can extract events with massive types is a long-standing target in Event Extraction (EE). In doing so, the challenge comes from…
Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery
Guankun Wang, Long Bai, Wan Jun Nah +7
Recent advancements in Surgical Visual Question Answering (Surgical-VQA) and related region grounding have shown great promise for robotic and medical applications, addressing the…
Transferring Knowledge from High-Quality to Low-Quality MRI for Adult Glioma Diagnosis
Yanguang Zhao, Long Bai, Zhaoxi Zhang +3
Glioma, a common and deadly brain tumor, requires early diagnosis for improved prognosis. However, low-quality Magnetic Resonance Imaging (MRI) technology in Sub-Saharan Africa (SS…
Approximation of Kolmogorov-Smirnov Test Statistics
Long Bai, David Kalaj
Motivated by the weak limit of the Kolmogorov-Smirnov test statistics, in this contribution, we concern the asymptotics of \begin{align*} \mathbb{P}\left\{\sup_{\boldsymbol{x}\in […
Benchmarking Pathology Foundation Models for Spatial Domain Understanding
Bokai Zhao, Yiyang Zhang, Yuanchi Zhu +6
Pathology foundation models (PFMs) have emerged as a core approach for learning transferable representations from whole slide images (WSIs), and they are typically benchmarked thro…
Towards Knowledgeable Deep Research: Framework and Benchmark
Wenxuan Liu, Zixuan Li, Long Bai +13
Deep Research (DR) requires LLM agents to autonomously perform multi-step information seeking, processing, and reasoning to generate comprehensive reports. In contrast to existing…
Rethinking Exemplars for Continual Semantic Segmentation in Endoscopy Scenes: Entropy-based Mini-Batch Pseudo-Replay
Guankun Wang, Long Bai, Yanan Wu +2
Endoscopy is a widely used technique for the early detection of diseases or robotic-assisted minimally invasive surgery (RMIS). Numerous deep learning (DL)-based research works hav…
SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian Splatting
Yiming Huang, Long Bai, Beilei Cui +6
In contemporary surgical research and practice, accurately comprehending 3D surgical scenes with text-promptable capabilities is particularly crucial for surgical planning and real…
GeoCFNet: Geometry-Aware Confidence Field Network for Robot-Assisted Endoscopic Submucosal Dissection
Rui Tang, Guankun Wang, Long Bai +5
Advanced surgical robotics has made robot-assisted endoscopic submucosal dissection (ESD) a promising approach for the en-bloc resection of large lesions, with the potential to red…
Advancing Dense Endoscopic Reconstruction with Gaussian Splatting-driven Surface Normal-aware Tracking and Mapping
Yiming Huang, Beilei Cui, Long Bai +5
Simultaneous Localization and Mapping (SLAM) is essential for precise surgical interventions and robotic tasks in minimally invasive procedures. While recent advancements in 3D Gau…
Extremes of multifractional Brownian motion
Long Bai
Let be the standard Multifractional Brownian Motion(mBm), in this contribution we are concerned with the exact asymptotics of \begin{eqnarra…
Retrieval-Augmented Code Generation for Universal Information Extraction
Yucan Guo, Zixuan Li, Xiaolong Jin +8
Information Extraction (IE) aims to extract structural knowledge (e.g., entities, relations, events) from natural language texts, which brings challenges to existing methods due to…
Joint Sparse Representations and Coupled Dictionary Learning in Multi-Source Heterogeneous Image Pseudo-color Fusion
Long Bai, Shilong Yao, Kun Gao +5
Considering that Coupled Dictionary Learning (CDL) method can obtain a reasonable linear mathematical relationship between resource images, we propose a novel CDL-based Synthetic A…
Recognizing Surgical Phases Anywhere: Few-Shot Test-time Adaptation and Task-graph Guided Refinement
Kun Yuan, Tingxuan Chen, Shi Li +9
The complexity and diversity of surgical workflows, driven by heterogeneous operating room settings, institutional protocols, and anatomical variability, present a significant chal…
ETSM: Automating Dissection Trajectory Suggestion and Confidence Map-Based Safety Margin Prediction for Robot-assisted Endoscopic Submucosal Dissection
Mengya Xu, Wenjin Mo, Guankun Wang +7
Robot-assisted Endoscopic Submucosal Dissection (ESD) improves the surgical procedure by providing a more comprehensive view through advanced robotic instruments and bimanual opera…
Estimation of Change-point Models
Long Bai
We consider the testing and estimation of change-points, locations where the distribution abruptly changes, in a sequence of observations. Motivated by this problem, in this contri…
Semantic Structure Enhanced Event Causality Identification
Zhilei Hu, Zixuan Li, Xiaolong Jin +4
Event Causality Identification (ECI) aims to identify causal relations between events in unstructured texts. This is a very challenging task, because causal relations are usually e…
Intuitive Surgical SurgToolLoc and SurgVU Challenges Results: 2022-2025
Aneeq Zia, Max Berniker, Rogerio Garcia Nespolo +153
Robotic assisted (RA) surgery promises to transform surgical intervention. Intuitive Surgical is committed to fostering these changes and the machine learning models and algorithms…
EndoDAC: Efficient Adapting Foundation Model for Self-Supervised Depth Estimation from Any Endoscopic Camera
Beilei Cui, Mobarakol Islam, Long Bai +2
Depth estimation plays a crucial role in various tasks within endoscopic surgery, including navigation, surface reconstruction, and augmented reality visualization. Despite the sig…
Privacy-Preserving Synthetic Continual Semantic Segmentation for Robotic Surgery
Mengya Xu, Mobarakol Islam, Long Bai +1
Deep Neural Networks (DNNs) based semantic segmentation of the robotic instruments and tissues can enhance the precision of surgical activities in robot-assisted surgery. However,…
Geometry OR Tracker: Universal Geometric Operating Room Tracking
Yihua Shao, Kang Chen, Feng Xue +6
In operating rooms (OR), world-scale multi-view 3D tracking supports downstream applications such as surgeon behavior recognition, where physically meaningful quantities such as di…
Domain Adaptive Sim-to-Real Segmentation of Oropharyngeal Organs Towards Robot-assisted Intubation
Guankun Wang, Tian-Ao Ren, Jiewen Lai +2
Robotic-assisted tracheal intubation requires the robot to distinguish anatomical features like an experienced physician using deep-learning techniques. However, real datasets of o…
EndoVLA: Dual-Phase Vision-Language-Action Model for Autonomous Tracking in Endoscopy
Chi Kit Ng, Long Bai, Guankun Wang +6
In endoscopic procedures, autonomous tracking of abnormal regions and following circumferential cutting markers can significantly reduce the cognitive burden on endoscopists. Howev…
Web-based Augmented Reality with Auto-Scaling and Real-Time Head Tracking towards Markerless Neurointerventional Preoperative Planning and Training of Head-mounted Robotic Needle Insertion
Hon Lung Ho, Yupeng Wang, An Wang +2
Neurosurgery requires exceptional precision and comprehensive preoperative planning to ensure optimal patient outcomes. Despite technological advancements, there remains a need for…
Beyond Dialogue Time: Temporal Semantic Memory for Personalized LLM Agents
Miao Su, Yucan Guo, Zhongni Hou +8
Memory enables Large Language Model (LLM) agents to perceive, store, and use information from past dialogues, which is essential for personalization. However, existing methods fail…
LighTDiff: Surgical Endoscopic Image Low-Light Enhancement with T-Diffusion
Tong Chen, Qingcheng Lyu, Long Bai +5
Advances in endoscopy use in surgeries face challenges like inadequate lighting. Deep learning, notably the Denoising Diffusion Probabilistic Model (DDPM), holds promise for low-li…
KnowCoder-V2: Deep Knowledge Analysis
Zixuan Li, Wenxuan Liu, Long Bai +13
Deep knowledge analysis tasks always involve the systematic extraction and association of knowledge from large volumes of data, followed by logical reasoning to discover insights.…
EndoUIC: Promptable Diffusion Transformer for Unified Illumination Correction in Capsule Endoscopy
Long Bai, Tong Chen, Qiaozhi Tan +10
Wireless Capsule Endoscopy (WCE) is highly valued for its non-invasive and painless approach, though its effectiveness is compromised by uneven illumination from hardware constrain…
Geo-RepNet: Geometry-Aware Representation Learning for Surgical Phase Recognition in Endoscopic Submucosal Dissection
Rui Tang, Haochen Yin, Guankun Wang +5
Surgical phase recognition plays a critical role in developing intelligent assistance systems for minimally invasive procedures such as Endoscopic Submucosal Dissection (ESD). Howe…
Surgical-VQLA: Transformer with Gated Vision-Language Embedding for Visual Question Localized-Answering in Robotic Surgery
Long Bai, Mobarakol Islam, Lalithkumar Seenivasan +1
Despite the availability of computer-aided simulators and recorded videos of surgical procedures, junior residents still heavily rely on experts to answer their queries. However, e…
Learning to Adapt Foundation Model DINOv2 for Capsule Endoscopy Diagnosis
Bowen Zhang, Ying Chen, Long Bai +5
Foundation models have become prominent in computer vision, achieving notable success in various tasks. However, their effectiveness largely depends on pre-training with extensive…
The Exploration and Evaluation of Generating Affective 360 Panoramic VR Environments Through Neural Style Transfer
Yanheng Li, Long Bai, Yaxuan Mao +4
Affective virtual reality (VR) environments with varying visual style can impact users' valence and arousal responses. We applied Neural Style Transfer (NST) to generate 360$^\circ…
EndoDDC: Learning Sparse to Dense Reconstruction for Endoscopic Robotic Navigation via Diffusion Depth Completion
Yinheng Lin, Yiming Huang, Beilei Cui +4
Accurate depth estimation plays a critical role in the navigation of endoscopic surgical robots, forming the foundation for 3D reconstruction and safe instrument guidance. Fine-tun…
Sample-adaptive Augmentation for Point Cloud Recognition Against Real-world Corruptions
Jie Wang, Lihe Ding, Tingfa Xu +4
Robust 3D perception under corruption has become an essential task for the realm of 3D vision. While current data augmentation techniques usually perform random transformations on…
KnowCoder-X: Boosting Multilingual Information Extraction via Code
Yuxin Zuo, Wenxuan Jiang, Wenxuan Liu +7
Empirical evidence indicates that LLMs exhibit spontaneous cross-lingual alignment. However, although LLMs show promising cross-lingual alignment in Information Extraction (IE), a…
Learning to Efficiently Adapt Foundation Models for Self-Supervised Endoscopic 3D Scene Reconstruction from Any Cameras
Beilei Cui, Long Bai, Mobarakol Islam +8
Accurate 3D scene reconstruction is essential for numerous medical tasks. Given the challenges in obtaining ground truth data, there has been an increasing focus on self-supervised…
More than Segmentation: Benchmarking SAM 3 for Segmentation, 3D Perception, and Reconstruction in Robotic Surgery
Wenzhen Dong, Jieming Yu, Yiming Huang +5
The recent SAM 3 and SAM 3D have introduced significant advancements over the predecessor, SAM 2, particularly with the integration of language-based segmentation and enhanced 3D p…
Extremes of Locally-stationary Chi-square processes on discrete grids
Long Bai
For centered Gaussian processes, the chi-square process appears naturally as limiting processes in various statistical…
Registering Neural 4D Gaussians for Endoscopic Surgery
Yiming Huang, Beilei Cui, Ikemura Kei +3
The recent advance in neural rendering has enabled the ability to reconstruct high-quality 4D scenes using neural networks. Although 4D neural reconstruction is popular, registrati…
EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery
Guankun Wang, Long Bai, Junyi Wang +13
Recently, Multimodal Large Language Models (MLLMs) have demonstrated their immense potential in computer-aided diagnosis and decision-making. In the context of robotic-assisted sur…
TR2M: Transferring Monocular Relative Depth to Metric Depth with Language Descriptions and Dual-Level Scale-Oriented Contrast
Beilei Cui, Yiming Huang, Long Bai +1
This work presents a generalizable framework to transfer relative depth to metric depth. Current monocular depth estimation methods are mainly divided into metric depth estimation…
CAT-ViL: Co-Attention Gated Vision-Language Embedding for Visual Question Localized-Answering in Robotic Surgery
Long Bai, Mobarakol Islam, Hongliang Ren
Medical students and junior surgeons often rely on senior surgeons and specialists to answer their questions when learning surgery. However, experts are often busy with clinical an…
Code-on-Graph: Iterative Programmatic Reasoning via Large Language Models on Knowledge Graphs
Weiwei Ding, Zixuan Li, Long Bai +7
Knowledge Graphs (KGs) are widely used to mitigate the limitations of Large Language Models (LLMs), such as outdated knowledge and hallucinations. Existing LLM-KG integration frame…
Revisiting Distillation for Continual Learning on Visual Question Localized-Answering in Robotic Surgery
Long Bai, Mobarakol Islam, Hongliang Ren
The visual-question localized-answering (VQLA) system can serve as a knowledgeable assistant in surgical education. Except for providing text-based answers, the VQLA system can hig…
GOPT: Generalizable Online 3D Bin Packing via Transformer-based Deep Reinforcement Learning
Heng Xiong, Changrong Guo, Jian Peng +5
Robotic object packing has broad practical applications in the logistics and automation industry, often formulated by researchers as the online 3D Bin Packing Problem (3D-BPP). How…
CapsDT: Diffusion-Transformer for Capsule Robot Manipulation
Xiting He, Mingwu Su, Xinqi Jiang +3
Vision-Language-Action (VLA) models have emerged as a prominent research area, showcasing significant potential across a variety of applications. However, their performance in endo…
Parameter-Efficient Adaptation of SAM3 for Prompt-Driven Surgical Concept Segmentation
Changjing Liu, Yiming Huang, Beilei Cui +5
Efficient surgical segmentation empowers clinical diagnosis, intraoperative monitoring, and downstream robotic pipelines for reconstruction and simulation. Although prompt-driven f…
KnowCoder: Coding Structured Knowledge into LLMs for Universal Information Extraction
Zixuan Li, Yutao Zeng, Yuxin Zuo +14
In this paper, we propose KnowCoder, a Large Language Model (LLM) to conduct Universal Information Extraction (UIE) via code generation. KnowCoder aims to develop a kind of unified…
Surgical-DINO: Adapter Learning of Foundation Models for Depth Estimation in Endoscopic Surgery
Beilei Cui, Mobarakol Islam, Long Bai +1
Purpose: Depth estimation in robotic surgery is vital in 3D reconstruction, surgical navigation and augmented reality visualization. Although the foundation model exhibits outstand…
NeuroABench: A Multimodal Evaluation Benchmark for Neurosurgical Anatomy Identification
Ziyang Song, Zelin Zang, Xiaofan Ye +7
Multimodal Large Language Models (MLLMs) have shown significant potential in surgical video understanding. With improved zero-shot performance and more effective human-machine inte…
HiSMatch: Historical Structure Matching based Temporal Knowledge Graph Reasoning
Zixuan Li, Zhongni Hou, Saiping Guan +7
A Temporal Knowledge Graph (TKG) is a sequence of KGs with respective timestamps, which adopts quadruples in the form of (\emph{subject}, \emph{relation}, \emph{object}, \emph{time…