5 papers
Bootstrapping Vision-Language Model for Hysteroscopic Surgical Scene Segmentation
Jun Huang, Meiyi Chen, Zijie Yue +6
Hysteroscopic surgical scene segmentation plays a pivotal role in understanding the hysteroscopic intraoperative environment as well as computer-assisted intervention. However, thi…
Weakly-Supervised Referring Video Object Segmentation through Text Supervision
Miaojing Shi, Jun Huang, Zijie Yue +1
Referring video object segmentation (RVOS) aims to segment the target instance in a video, referred by a text expression. Conventional approaches are mostly supervised learning, re…
Surg-R1: A Hierarchical Reasoning Foundation Model for Scalable and Interpretable Surgical Decision Support with Multi-Center Clinical Validation
Jian Jiang, Chenxi Lin, Yiming Gu +24
Surgical scene understanding demands not only accurate predictions but also interpretable reasoning that surgeons can verify against clinical expertise. However, existing surgical…
Bootstrapping MLLM for Weakly-Supervised Class-Agnostic Object Counting
Xiaowen Zhang, Zijie Yue, Yong Luo +3
Object counting is a fundamental task in computer vision, with broad applicability in many real-world scenarios. Fully-supervised counting methods require costly point-level annota…
Development and validation of an AI foundation model for endoscopic diagnosis of esophagogastric junction adenocarcinoma: a cohort and deep learning study
Yikun Ma, Bo Li, Ying Chen +20
The early detection of esophagogastric junction adenocarcinoma (EGJA) is crucial for improving patient prognosis, yet its current diagnosis is highly operator-dependent. This paper…