papers

Publications (116)

cs.CV2026

SegSTRONG-C: Segmenting Surgical Tools Robustly On Non-adversarial Generated Corruptions -- An EndoVis'24 Challenge

Hao Ding, Yuqian Zhang, Tuxun Lu +39

Surgical data science has seen rapid advancement with the excellent performance of end-to-end deep neural networks (DNNs). Despite their successes, DNNs have been proven susceptibl…

cs.LG2022

What is Event Knowledge Graph: A Survey

Saiping Guan, Xueqi Cheng, Long Bai +5

Besides entity-centric knowledge, usually organized as Knowledge Graph (KG), events are also an essential kind of knowledge in the world, which trigger the spring up of event-centr…

math.PR2018

Extremes of Gaussian chaos processes with Trend

Long Bai

Let be a Gaussian vector process and let be a continuous homogeneous functi…

cs.CV2024

CoPESD: A Multi-Level Surgical Motion Dataset for Training Large Vision-Language Models to Co-Pilot Endoscopic Submucosal Dissection

Guankun Wang, Han Xiao, Huxin Gao +6

submucosal dissection (ESD) enables rapid resection of large lesions, minimizing recurrence rates and improving long-term overall survival. Despite these advantages, ESD is technic…

cs.CL2024

Nested Event Extraction upon Pivot Element Recogniton

Weicheng Ren, Zixuan Li, Xiaolong Jin +6

Nested Event Extraction (NEE) aims to extract complex event structures where an event contains other events as its arguments recursively. Nested events involve a kind of Pivot Elem…

cs.CL2025

Towards Robust Universal Information Extraction: Benchmark, Evaluation, and Solution

Jizhao Zhu, Akang Shi, Zixuan Li +4

In this paper, we aim to enhance the robustness of Universal Information Extraction (UIE) by introducing a new benchmark dataset, a comprehensive evaluation, and a feasible solutio…

eess.IV2021

The Influence of Age and Gender Information on the Diagnosis of Diabetic Retinopathy: Based on Neural Networks

Long Bai, Sihang Chen, Mingyang Gao +3

This paper proposes the importance of age and gender information in the diagnosis of diabetic retinopathy. We utilized Deep Residual Neural Networks (ResNet) and Densely Connected…

math.PR2018

Extremes of -norm of Vector-valued Gaussian processes with Trend

Long Bai

Let be a Gaussian vector process and be a continuous function. The asymptotics of distribution of $\left\|\boldsymbol{X}(t)\right\…

eess.IV2026

DaX: Learning General Pathology Representations Across Scales

Bokai Zhao, Yiyang Zhang, Long Bai +3

Computational pathology requires visual representations that transfer across diverse clinical endpoints and remain robust to variation in magnification, staining, scanner type, sli…

cs.RO2026

GeoLanG: Geometry-Aware Language-Guided Grasping with Unified RGB-D Multimodal Learning

Rui Tang, Guankun Wang, Long Bai +6

Language-guided grasping has emerged as a promising paradigm for enabling robots to identify and manipulate target objects through natural language instructions, yet it remains hig…

cs.RO2026

TMR-VLA:Vision-Language-Action Model for Magnetic Motion Control of Tri-leg Silicone-based Soft Robot

Ruijie Tang, Chi Kit Ng, Kaixuan Wu +5

In-vivo environments, magnetically actuated soft robots offer advantages such as wireless operation and precise control, showing promising potential for painless detection and ther…

math.PR2017

Parisian ruin of Brownian motion risk model over an infinite-time horizon

Long Bai

Let be a standard Brownian motion. In this paper, we derive the exact asymptotics of the probability of Parisian ruin on infinite time horizon for the follo…

cs.AI2026

FinDeepIndicator: Benchmarking Deep Research Agents in End-to-End Financial Indicator Construction

Chaoqun Yang, Fengbin Zhu, Xinyu Lin +5

Financial indicators are essential tools for transforming raw financial data into interpretable measures for various downstream tasks, such as valuation, risk assessment, and econo…

cs.CV2026

SurgOnAir: Hierarchy-Aware Real-Time Surgical Video Commentary

Jingyi He, Yue Zhou, Long Bai +3

Understanding surgical workflow in real time is fundamental for intelligent surgical embodiment, where AI systems continuously perceive and respond as surgery proceeds. In the oper…

cs.CL2022

Rich Event Modeling for Script Event Prediction

Long Bai, Saiping Guan, Zixuan Li +3

Script is a kind of structured knowledge extracted from texts, which contains a sequence of events. Based on such knowledge, script event prediction aims to predict the subsequent…

cs.CL2026

Event Ontology Expansion via LLM-Based Conceptualization

Weicheng Ren, Zixuan Li, Long Bai +3

Event ontology expansion aims to discover emerging event types from data and extend them to appropriate positions in the existing event ontology.. Existing methods typically cluste…

cs.AI2026

Surg-UniWorld: A Unified Surgical World Model with Multimodal Control Experts

Rulin Zhou, Wanhao Liu, Guoheng Ma +8

Controllable surgical world models can provide a generative foundation for surgical artificial intelligence and simulation by synthesizing realistic instrument--tissue interactions…

cs.CV2024

Surgical-VQLA++: Adversarial Contrastive Learning for Calibrated Robust Visual Question-Localized Answering in Robotic Surgery

Long Bai, Guankun Wang, Mobarakol Islam +3

Medical visual question answering (VQA) bridges the gap between visual information and clinical decision-making, enabling doctors to extract understanding from clinical images and…

cs.CL2021

Integrating Deep Event-Level and Script-Level Information for Script Event Prediction

Long Bai, Saiping Guan, Jiafeng Guo +3

Scripts are structured sequences of events together with the participants, which are extracted from the texts.Script event prediction aims to predict the subsequent event given the…

cs.CV2025

Endo-4DGX: Robust Endoscopic Scene Reconstruction and Illumination Correction with Gaussian Splatting

Yiming Huang, Long Bai, Beilei Cui +7

Accurate reconstruction of soft tissue is crucial for advancing automation in image-guided robotic surgery. The recent 3D Gaussian Splatting (3DGS) techniques and their variants, 4…

math.PR2018

Extremes of vector-valued Gaussian processes with Trend

Long Bai, Krzysztof Debicki, Peng Liu

Let be a centered vector-valued Gaussian process with independent components and continuous trajectories, and $h…

cs.CV2025

SurgPLAN++: Universal Surgical Phase Localization Network for Online and Offline Inference

Zhen Chen, Xingjian Luo, Jinlin Wu +5

Surgical phase recognition is critical for assisting surgeons in understanding surgical videos. Existing studies focused more on online surgical phase recognition, by leveraging pr…

math.PR2021

Extremes of Gaussian random fields with non-additive dependence structure

Long Bai, Krzysztof Debicki, Peng Liu

We derive exact asymptotics of for a centered Gaussian field $X(\mathbf{t}),~…

cs.CV2024

EndoOOD: Uncertainty-aware Out-of-distribution Detection in Capsule Endoscopy Diagnosis

Qiaozhi Tan, Long Bai, Guankun Wang +2

Wireless capsule endoscopy (WCE) is a non-invasive diagnostic procedure that enables visualization of the gastrointestinal (GI) tract. Deep learning-based methods have shown effect…

math.PR2016

Extremes of -Locally stationary Gaussian processes with non-constant variances

Long Bai

With motivation from K. DÈ©bicki and P. Kisowski (2007), in this paper we derive the exact tail asymptotics of -locally stationary Gaussian processes with non-constant varia…

cs.CV2025

PvNeXt: Rethinking Network Design and Temporal Motion for Point Cloud Video Recognition

Jie Wang, Tingfa Xu, Lihe Ding +3

Point cloud video perception has become an essential task for the realm of 3D vision. Current 4D representation learning techniques typically engage in iterative processing coupled…

cs.CV2023

Semi-supervised Learning for Segmentation of Bleeding Regions in Video Capsule Endoscopy

Hechen Li, Yanan Wu, Long Bai +3

In the realm of modern diagnostic technology, video capsule endoscopy (VCE) is a standout for its high efficacy and non-invasive nature in diagnosing various gastrointestinal (GI)…

cs.CV2026

Where It Moves, It Matters: Referring Surgical Instrument Segmentation via Motion

Meng Wei, Kun Yuan, Shi Li +7

Enabling intuitive, language-driven interaction with surgical scenes is a critical step toward intelligent operating rooms and autonomous surgical robotic assistance. However, the…

cs.CV2024

ASI-Seg: Audio-Driven Surgical Instrument Segmentation with Surgeon Intention Understanding

Zhen Chen, Zongming Zhang, Wenwu Guo +5

Surgical instrument segmentation is crucial in surgical scene understanding, thereby facilitating surgical safety. Existing algorithms directly detected all instruments of pre-defi…

math.ST2019

On generalized Piterbarg-Berman function

Chengxiu Ling, Hong Zhang, Long Bai

This paper aims to evaluate the Piterbarg-Berman function given by $$\mathcal{P\!B}_α^h(x, E) = \int_\mathbb{R}e^z\mathbb{P} \left\{{\int_E \mathbb{I}\left(\sqrt2B_α(t) - |t|^α-…

cs.CV2026

Does DINOv3 Set a New Medical Vision Standard? Benchmarking 2D and 3D Classification, Segmentation, and Registration

Che Liu, Yinda Chen, Haoyuan Shi +21

The advent of large-scale vision foundation models, pre-trained on diverse natural images, has marked a paradigm shift in computer vision. However, how the frontier vision foundati…

math.PR2018

Drawdown and drawup for fractional Brownian motion with trend

Long Bai, Peng Liu

In this paper, we consider the drawdown and drawup of the fractional Brownian motion with trend, which corresponds to the logarithm of geometric fractional Brownian motion represen…

cs.CV2024

Endo-4DGS: Endoscopic Monocular Scene Reconstruction with 4D Gaussian Splatting

Yiming Huang, Beilei Cui, Long Bai +4

In the realm of robot-assisted minimally invasive surgery, dynamic scene reconstruction can significantly enhance downstream tasks and improve surgical outcomes. Neural Radiance Fi…

cs.AI2023

Domain Adaptive Sim-to-Real Segmentation of Oropharyngeal Organs

Guankun Wang, Tian-Ao Ren, Jiewen Lai +2

Video-assisted transoral tracheal intubation (TI) necessitates using an endoscope that helps the physician insert a tracheal tube into the glottis instead of the esophagus. The gro…

math.PR2018

Ruin problem for Brownian motion risk model with interest rate and tax payment

Long Bai, Peng Liu

Let be a Brownian motion. Consider the Brownian motion risk model with interest rate collection and tax payment defined by \begin{align}\label{Rudef} \widetilde{…

cs.CV2025

Can DeepSeek Reason Like a Surgeon? An Empirical Evaluation for Vision-Language Understanding in Robotic-Assisted Surgery

Boyi Ma, Yanguang Zhao, Jie Wang +5

The DeepSeek models have shown exceptional performance in general scene understanding, question-answering (QA), and text generation tasks, owing to their efficient training paradig…

cs.CL2024

Class-Incremental Few-Shot Event Detection

Kailin Zhao, Xiaolong Jin, Long Bai +2

Event detection is one of the fundamental tasks in information extraction and knowledge graph. However, a realistic event detection system often needs to deal with new event classe…

cs.CV2024

SAM 2 in Robotic Surgery: An Empirical Evaluation for Robustness and Generalization in Surgical Video Segmentation

Jieming Yu, An Wang, Wenzhen Dong +5

The recent Segment Anything Model (SAM) 2 has demonstrated remarkable foundational competence in semantic segmentation, with its memory mechanism and mask decoder further addressin…

cs.CL2024

An In-Context Schema Understanding Method for Knowledge Base Question Answering

Yantao Liu, Zixuan Li, Xiaolong Jin +5

The Knowledge Base Question Answering (KBQA) task aims to answer natural language questions based on a given knowledge base. Recently, Large Language Models (LLMs) have shown stron…

cs.CV2025

Multimodal Graph Representation Learning for Robust Surgical Workflow Recognition with Adversarial Feature Disentanglement

Long Bai, Boyi Ma, Ruohan Wang +8

Surgical workflow recognition is vital for automating tasks, supporting decision-making, and training novice surgeons, ultimately improving patient safety and standardizing procedu…

cs.RO2026

How can reasoning capability empower the AI copilot robot in endoscopic surgery

Guankun Wang, Long Bai, Hongliang Ren

Reasoning capability has significantly advanced complex logical inference and robotic decision-making in general domains. However, its potential in the Artificial Intelligence (AI)…

cs.CV2024

V-SfMLearner: Learning Monocular Depth and Ego-motion for Multimodal Wireless Capsule Endoscopy

Long Bai, Beilei Cui, Liangyu Wang +9

Deep learning can predict depth maps and capsule ego-motion from capsule endoscopy videos, aiding in 3D scene reconstruction and lesion localization. However, the collisions of the…

eess.IV2023

LLCaps: Learning to Illuminate Low-Light Capsule Endoscopy with Curved Wavelet Attention and Reverse Diffusion

Long Bai, Tong Chen, Yanan Wu +3

Wireless capsule endoscopy (WCE) is a painless and non-invasive diagnostic tool for gastrointestinal (GI) diseases. However, due to GI anatomical constraints and hardware manufactu…

cs.CV2024

A Review of 3D Reconstruction Techniques for Deformable Tissues in Robotic Surgery

Mengya Xu, Ziqi Guo, An Wang +2

As a crucial and intricate task in robotic minimally invasive surgery, reconstructing surgical scenes using stereo or monocular endoscopic video holds immense potential for clinica…

math.PR2016

Parisian Ruin of the Brownian Motion Risk Model with Constant Force of Interest

Long Bai, Li Luo

Let be a standard Brownian motion. Define a risk process \label{Rudef} R_u^δ(t)=e^{δt}\left(u+c\int^{t}_{0}e^{-δs}d s-σ\int_{0}^{t}e^{-δs}d B(s)\right)…

cs.CV2026

Comparative validation of surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation in endoscopy: Results of the PhaKIR 2024 challenge

Tobias Rueckert, David Rauber, Raphaela Maerkl +58

Reliable recognition and localization of surgical instruments in endoscopic video recordings are foundational for a wide range of applications in computer- and robot-assisted minim…

cs.CV2024

SAR-RARP50: Segmentation of surgical instrumentation and Action Recognition on Robot-Assisted Radical Prostatectomy Challenge

Dimitrios Psychogyios, Emanuele Colleoni, Beatrice Van Amsterdam +47

Surgical tool segmentation and action recognition are fundamental building blocks in many computer-assisted intervention applications, ranging from surgical skills assessment to de…

cs.CV2026

EndoGSim: Physics-Aware 4D Dynamic Endoscopic Scene Simulations via MLLM-Guided Gaussian Splatting

Changjing Liu, Yiming Huang, Long Bai +2

In robot-assisted minimally invasive surgery, high-fidelity dynamic endoscopic scene reconstruction and simulation are crucial to enhancing downstream tasks and advancing surgical…

cs.CV2025

SurgSora: Object-Aware Diffusion Model for Controllable Surgical Video Generation

Tong Chen, Shuya Yang, Junyi Wang +3

Surgical video generation can enhance medical education and research, but existing methods lack fine-grained motion control and realism. We introduce SurgSora, a framework that gen…

cs.CL2025

Towards Event Extraction with Massive Types: LLM-based Collaborative Annotation and Partitioning Extraction

Wenxuan Liu, Zixuan Li, Long Bai +5

Developing a general-purpose extraction system that can extract events with massive types is a long-standing target in Event Extraction (EE). In doing so, the challenge comes from…

cs.CV2025

Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery

Guankun Wang, Long Bai, Wan Jun Nah +7

Recent advancements in Surgical Visual Question Answering (Surgical-VQA) and related region grounding have shown great promise for robotic and medical applications, addressing the…

eess.IV2024

Transferring Knowledge from High-Quality to Low-Quality MRI for Adult Glioma Diagnosis

Yanguang Zhao, Long Bai, Zhaoxi Zhang +3

Glioma, a common and deadly brain tumor, requires early diagnosis for improved prognosis. However, low-quality Magnetic Resonance Imaging (MRI) technology in Sub-Saharan Africa (SS…

math.PR2018

Approximation of Kolmogorov-Smirnov Test Statistics

Long Bai, David Kalaj

Motivated by the weak limit of the Kolmogorov-Smirnov test statistics, in this contribution, we concern the asymptotics of \begin{align*} \mathbb{P}\left\{\sup_{\boldsymbol{x}\in […

cs.CV2026

Benchmarking Pathology Foundation Models for Spatial Domain Understanding

Bokai Zhao, Yiyang Zhang, Yuanchi Zhu +6

Pathology foundation models (PFMs) have emerged as a core approach for learning transferable representations from whole slide images (WSIs), and they are typically benchmarked thro…

cs.AI2026

Towards Knowledgeable Deep Research: Framework and Benchmark

Wenxuan Liu, Zixuan Li, Long Bai +13

Deep Research (DR) requires LLM agents to autonomously perform multi-step information seeking, processing, and reasoning to generate comprehensive reports. In contrast to existing…

cs.CV2023

Rethinking Exemplars for Continual Semantic Segmentation in Endoscopy Scenes: Entropy-based Mini-Batch Pseudo-Replay

Guankun Wang, Long Bai, Yanan Wu +2

Endoscopy is a widely used technique for the early detection of diseases or robotic-assisted minimally invasive surgery (RMIS). Numerous deep learning (DL)-based research works hav…

eess.IV2025

SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian Splatting

Yiming Huang, Long Bai, Beilei Cui +6

In contemporary surgical research and practice, accurately comprehending 3D surgical scenes with text-promptable capabilities is particularly crucial for surgical planning and real…

cs.CV2026

GeoCFNet: Geometry-Aware Confidence Field Network for Robot-Assisted Endoscopic Submucosal Dissection

Rui Tang, Guankun Wang, Long Bai +5

Advanced surgical robotics has made robot-assisted endoscopic submucosal dissection (ESD) a promising approach for the en-bloc resection of large lesions, with the potential to red…

cs.CV2025

Advancing Dense Endoscopic Reconstruction with Gaussian Splatting-driven Surface Normal-aware Tracking and Mapping

Yiming Huang, Beilei Cui, Long Bai +5

Simultaneous Localization and Mapping (SLAM) is essential for precise surgical interventions and robotic tasks in minimally invasive procedures. While recent advancements in 3D Gau…

math.PR2019

Extremes of multifractional Brownian motion

Long Bai

Let be the standard Multifractional Brownian Motion(mBm), in this contribution we are concerned with the exact asymptotics of \begin{eqnarra…

cs.AI2023

Retrieval-Augmented Code Generation for Universal Information Extraction

Yucan Guo, Zixuan Li, Xiaolong Jin +8

Information Extraction (IE) aims to extract structural knowledge (e.g., entities, relations, events) from natural language texts, which brings challenges to existing methods due to…

eess.IV2023

Joint Sparse Representations and Coupled Dictionary Learning in Multi-Source Heterogeneous Image Pseudo-color Fusion

Long Bai, Shilong Yao, Kun Gao +5

Considering that Coupled Dictionary Learning (CDL) method can obtain a reasonable linear mathematical relationship between resource images, we propose a novel CDL-based Synthetic A…

cs.CV2025

Recognizing Surgical Phases Anywhere: Few-Shot Test-time Adaptation and Task-graph Guided Refinement

Kun Yuan, Tingxuan Chen, Shi Li +9

The complexity and diversity of surgical workflows, driven by heterogeneous operating room settings, institutional protocols, and anatomical variability, present a significant chal…

cs.RO2024

ETSM: Automating Dissection Trajectory Suggestion and Confidence Map-Based Safety Margin Prediction for Robot-assisted Endoscopic Submucosal Dissection

Mengya Xu, Wenjin Mo, Guankun Wang +7

Robot-assisted Endoscopic Submucosal Dissection (ESD) improves the surgical procedure by providing a more comprehensive view through advanced robotic instruments and bimanual opera…

math.PR2018

Estimation of Change-point Models

Long Bai

We consider the testing and estimation of change-points, locations where the distribution abruptly changes, in a sequence of observations. Motivated by this problem, in this contri…

cs.CL2023

Semantic Structure Enhanced Event Causality Identification

Zhilei Hu, Zixuan Li, Xiaolong Jin +4

Event Causality Identification (ECI) aims to identify causal relations between events in unstructured texts. This is a very challenging task, because causal relations are usually e…

cs.CV2026

Intuitive Surgical SurgToolLoc and SurgVU Challenges Results: 2022-2025

Aneeq Zia, Max Berniker, Rogerio Garcia Nespolo +153

Robotic assisted (RA) surgery promises to transform surgical intervention. Intuitive Surgical is committed to fostering these changes and the machine learning models and algorithms…

eess.IV2024

EndoDAC: Efficient Adapting Foundation Model for Self-Supervised Depth Estimation from Any Endoscopic Camera

Beilei Cui, Mobarakol Islam, Long Bai +2

Depth estimation plays a crucial role in various tasks within endoscopic surgery, including navigation, surface reconstruction, and augmented reality visualization. Despite the sig…

cs.CV2024

Privacy-Preserving Synthetic Continual Semantic Segmentation for Robotic Surgery

Mengya Xu, Mobarakol Islam, Long Bai +1

Deep Neural Networks (DNNs) based semantic segmentation of the robotic instruments and tissues can enhance the precision of surgical activities in robot-assisted surgery. However,…

cs.CV2026

Geometry OR Tracker: Universal Geometric Operating Room Tracking

Yihua Shao, Kang Chen, Feng Xue +6

In operating rooms (OR), world-scale multi-view 3D tracking supports downstream applications such as surgeon behavior recognition, where physically meaningful quantities such as di…

eess.IV2023

Domain Adaptive Sim-to-Real Segmentation of Oropharyngeal Organs Towards Robot-assisted Intubation

Guankun Wang, Tian-Ao Ren, Jiewen Lai +2

Robotic-assisted tracheal intubation requires the robot to distinguish anatomical features like an experienced physician using deep-learning techniques. However, real datasets of o…

cs.RO2025

EndoVLA: Dual-Phase Vision-Language-Action Model for Autonomous Tracking in Endoscopy

Chi Kit Ng, Long Bai, Guankun Wang +6

In endoscopic procedures, autonomous tracking of abnormal regions and following circumferential cutting markers can significantly reduce the cognitive burden on endoscopists. Howev…

cs.RO2024

Web-based Augmented Reality with Auto-Scaling and Real-Time Head Tracking towards Markerless Neurointerventional Preoperative Planning and Training of Head-mounted Robotic Needle Insertion

Hon Lung Ho, Yupeng Wang, An Wang +2

Neurosurgery requires exceptional precision and comprehensive preoperative planning to ensure optimal patient outcomes. Despite technological advancements, there remains a need for…

cs.AI2026

Beyond Dialogue Time: Temporal Semantic Memory for Personalized LLM Agents

Miao Su, Yucan Guo, Zhongni Hou +8

Memory enables Large Language Model (LLM) agents to perceive, store, and use information from past dialogues, which is essential for personalization. However, existing methods fail…

eess.IV2024

LighTDiff: Surgical Endoscopic Image Low-Light Enhancement with T-Diffusion

Tong Chen, Qingcheng Lyu, Long Bai +5

Advances in endoscopy use in surgeries face challenges like inadequate lighting. Deep learning, notably the Denoising Diffusion Probabilistic Model (DDPM), holds promise for low-li…

cs.AI2025

KnowCoder-V2: Deep Knowledge Analysis

Zixuan Li, Wenxuan Liu, Long Bai +13

Deep knowledge analysis tasks always involve the systematic extraction and association of knowledge from large volumes of data, followed by logical reasoning to discover insights.…

eess.IV2024

EndoUIC: Promptable Diffusion Transformer for Unified Illumination Correction in Capsule Endoscopy

Long Bai, Tong Chen, Qiaozhi Tan +10

Wireless Capsule Endoscopy (WCE) is highly valued for its non-invasive and painless approach, though its effectiveness is compromised by uneven illumination from hardware constrain…

cs.CV2025

Geo-RepNet: Geometry-Aware Representation Learning for Surgical Phase Recognition in Endoscopic Submucosal Dissection

Rui Tang, Haochen Yin, Guankun Wang +5

Surgical phase recognition plays a critical role in developing intelligent assistance systems for minimally invasive procedures such as Endoscopic Submucosal Dissection (ESD). Howe…

cs.CV2023

Surgical-VQLA: Transformer with Gated Vision-Language Embedding for Visual Question Localized-Answering in Robotic Surgery

Long Bai, Mobarakol Islam, Lalithkumar Seenivasan +1

Despite the availability of computer-aided simulators and recorded videos of surgical procedures, junior residents still heavily rely on experts to answer their queries. However, e…

cs.CV2024

Learning to Adapt Foundation Model DINOv2 for Capsule Endoscopy Diagnosis

Bowen Zhang, Ying Chen, Long Bai +5

Foundation models have become prominent in computer vision, achieving notable success in various tasks. However, their effectiveness largely depends on pre-training with extensive…

cs.HC2023

The Exploration and Evaluation of Generating Affective 360 Panoramic VR Environments Through Neural Style Transfer

Yanheng Li, Long Bai, Yaxuan Mao +4

Affective virtual reality (VR) environments with varying visual style can impact users' valence and arousal responses. We applied Neural Style Transfer (NST) to generate 360$^\circ…

cs.CV2026

EndoDDC: Learning Sparse to Dense Reconstruction for Endoscopic Robotic Navigation via Diffusion Depth Completion

Yinheng Lin, Yiming Huang, Beilei Cui +4

Accurate depth estimation plays a critical role in the navigation of endoscopic surgical robots, forming the foundation for 3D reconstruction and safe instrument guidance. Fine-tun…

cs.CV2023

Sample-adaptive Augmentation for Point Cloud Recognition Against Real-world Corruptions

Jie Wang, Lihe Ding, Tingfa Xu +4

Robust 3D perception under corruption has become an essential task for the realm of 3D vision. While current data augmentation techniques usually perform random transformations on…

cs.CL2025

KnowCoder-X: Boosting Multilingual Information Extraction via Code

Yuxin Zuo, Wenxuan Jiang, Wenxuan Liu +7

Empirical evidence indicates that LLMs exhibit spontaneous cross-lingual alignment. However, although LLMs show promising cross-lingual alignment in Information Extraction (IE), a…

cs.CV2025

Learning to Efficiently Adapt Foundation Models for Self-Supervised Endoscopic 3D Scene Reconstruction from Any Cameras

Beilei Cui, Long Bai, Mobarakol Islam +8

Accurate 3D scene reconstruction is essential for numerous medical tasks. Given the challenges in obtaining ground truth data, there has been an increasing focus on self-supervised…

cs.CV2025

More than Segmentation: Benchmarking SAM 3 for Segmentation, 3D Perception, and Reconstruction in Robotic Surgery

Wenzhen Dong, Jieming Yu, Yiming Huang +5

The recent SAM 3 and SAM 3D have introduced significant advancements over the predecessor, SAM 2, particularly with the integration of language-based segmentation and enhanced 3D p…

math.PR2018

Extremes of Locally-stationary Chi-square processes on discrete grids

Long Bai

For centered Gaussian processes, the chi-square process appears naturally as limiting processes in various statistical…

cs.RO2024

Registering Neural 4D Gaussians for Endoscopic Surgery

Yiming Huang, Beilei Cui, Ikemura Kei +3

The recent advance in neural rendering has enabled the ability to reconstruct high-quality 4D scenes using neural networks. Although 4D neural reconstruction is popular, registrati…

cs.CV2025

EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery

Guankun Wang, Long Bai, Junyi Wang +13

Recently, Multimodal Large Language Models (MLLMs) have demonstrated their immense potential in computer-aided diagnosis and decision-making. In the context of robotic-assisted sur…

cs.CV2026

TR2M: Transferring Monocular Relative Depth to Metric Depth with Language Descriptions and Dual-Level Scale-Oriented Contrast

Beilei Cui, Yiming Huang, Long Bai +1

This work presents a generalizable framework to transfer relative depth to metric depth. Current monocular depth estimation methods are mainly divided into metric depth estimation…

cs.CV2023

CAT-ViL: Co-Attention Gated Vision-Language Embedding for Visual Question Localized-Answering in Robotic Surgery

Long Bai, Mobarakol Islam, Hongliang Ren

Medical students and junior surgeons often rely on senior surgeons and specialists to answer their questions when learning surgery. However, experts are often busy with clinical an…

cs.AI2026

Code-on-Graph: Iterative Programmatic Reasoning via Large Language Models on Knowledge Graphs

Weiwei Ding, Zixuan Li, Long Bai +7

Knowledge Graphs (KGs) are widely used to mitigate the limitations of Large Language Models (LLMs), such as outdated knowledge and hallucinations. Existing LLM-KG integration frame…

cs.CV2023

Revisiting Distillation for Continual Learning on Visual Question Localized-Answering in Robotic Surgery

Long Bai, Mobarakol Islam, Hongliang Ren

The visual-question localized-answering (VQLA) system can serve as a knowledgeable assistant in surgical education. Except for providing text-based answers, the VQLA system can hig…

cs.RO2024

GOPT: Generalizable Online 3D Bin Packing via Transformer-based Deep Reinforcement Learning

Heng Xiong, Changrong Guo, Jian Peng +5

Robotic object packing has broad practical applications in the logistics and automation industry, often formulated by researchers as the online 3D Bin Packing Problem (3D-BPP). How…

cs.RO2025

CapsDT: Diffusion-Transformer for Capsule Robot Manipulation

Xiting He, Mingwu Su, Xinqi Jiang +3

Vision-Language-Action (VLA) models have emerged as a prominent research area, showcasing significant potential across a variety of applications. However, their performance in endo…

cs.CV2026

Parameter-Efficient Adaptation of SAM3 for Prompt-Driven Surgical Concept Segmentation

Changjing Liu, Yiming Huang, Beilei Cui +5

Efficient surgical segmentation empowers clinical diagnosis, intraoperative monitoring, and downstream robotic pipelines for reconstruction and simulation. Although prompt-driven f…

cs.LG2024

KnowCoder: Coding Structured Knowledge into LLMs for Universal Information Extraction

Zixuan Li, Yutao Zeng, Yuxin Zuo +14

In this paper, we propose KnowCoder, a Large Language Model (LLM) to conduct Universal Information Extraction (UIE) via code generation. KnowCoder aims to develop a kind of unified…

cs.CV2024

Surgical-DINO: Adapter Learning of Foundation Models for Depth Estimation in Endoscopic Surgery

Beilei Cui, Mobarakol Islam, Long Bai +1

Purpose: Depth estimation in robotic surgery is vital in 3D reconstruction, surgical navigation and augmented reality visualization. Although the foundation model exhibits outstand…

cs.CV2025

NeuroABench: A Multimodal Evaluation Benchmark for Neurosurgical Anatomy Identification

Ziyang Song, Zelin Zang, Xiaofan Ye +7

Multimodal Large Language Models (MLLMs) have shown significant potential in surgical video understanding. With improved zero-shot performance and more effective human-machine inte…

cs.AI2022

HiSMatch: Historical Structure Matching based Temporal Knowledge Graph Reasoning

Zixuan Li, Zhongni Hou, Saiping Guan +7

A Temporal Knowledge Graph (TKG) is a sequence of KGs with respective timestamps, which adopts quadruples in the form of (\emph{subject}, \emph{relation}, \emph{object}, \emph{time…