papers

Publications (121)

cs.CV2025

GMAI-VL-R1: Harnessing Reinforcement Learning for Multimodal Medical Reasoning

Yanzhou Su, Tianbin Li, Jiyao Liu +15

Recent advances in general medical AI have made significant strides, but existing models often lack the reasoning capabilities needed for complex medical decision-making. This pape…

eess.IV2024

GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI

Pengcheng Chen, Jin Ye, Guoan Wang +15

Large Vision-Language Models (LVLMs) are capable of handling diverse data types such as imaging, text, and physiological signals, and can be applied in various fields. In the medic…

cs.AI2026

A First-Principles Derivation of LLM Policy Optimization: From Expected Reward to GRPO and Its Structural Extensions

Jianghan Shen, Siqi Luo, Yue Li +9

Policy gradient algorithms for language models optimize the same objective , which has exactly two factors: the trajectory probability…

cs.CV2026

UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation

Yibin Wang, Zhimin Li, Yuhang Zang +8

Recent progress in text-to-image (T2I) generation underscores the importance of reliable benchmarks in evaluating how accurately generated images reflect the semantics of their tex…

cs.CV2025

Robust Multimodal Learning for Ophthalmic Disease Grading via Disentangled Representation

Xinkun Wang, Yifang Wang, Senwei Liang +7

This paper discusses how ophthalmologists often rely on multimodal data to improve diagnostic accuracy. However, complete multimodal data is rare in real-world applications due to…

eess.IV2022

Exploring Vanilla U-Net for Lesion Segmentation from Whole-body FDG-PET/CT Scans

Jin Ye, Haoyu Wang, Ziyan Huang +8

Tumor lesion segmentation is one of the most important tasks in medical image analysis. In clinical practice, Fluorodeoxyglucose Positron-Emission Tomography~(FDG-PET) is a widely…

cs.CV2026

MedQ-Deg: A Multidimensional Benchmark for Evaluating MLLMs Across Medical Image Quality Degradations

Jiyao Liu, Junzhi Ning, Chenglong Ma +14

Despite impressive performance on standard benchmarks, multimodal large language models (MLLMs) face critical challenges in real-world clinical environments where medical images in…

cs.CV2020

Attention-Driven Dynamic Graph Convolutional Network for Multi-Label Image Recognition

Jin Ye, Junjun He, Xiaojiang Peng +2

Recent studies often exploit Graph Convolutional Network (GCN) to model label dependencies to improve recognition accuracy for multi-label image recognition. However, constructing…

cs.CL2025

SciReasoner: Laying the Scientific Reasoning Ground Across Disciplines

Yizhou Wang, Chen Tang, Han Deng +29

We present a scientific reasoning foundation model that aligns natural language with heterogeneous scientific representations. The model is pretrained on a 206B-token corpus spanni…

cs.CL2025

MedAide: Information Fusion and Anatomy of Medical Intents via LLM-based Agent Collaboration

Dingkang Yang, Jinjie Wei, Mingcheng Li +8

In healthcare intelligence, the ability to fuse heterogeneous, multi-intent information from diverse clinical sources is fundamental to building reliable decision-making systems. L…

eess.IV2023

SA-Med2D-20M Dataset: Segment Anything in 2D Medical Imaging with 20 Million masks

Jin Ye, Junlong Cheng, Jianpin Chen +12

Segment Anything Model (SAM) has achieved impressive results for natural image segmentation with input prompts such as points and bounding boxes. Its success largely owes to massiv…

eess.IV2023

SegRap2023: A Benchmark of Organs-at-Risk and Gross Tumor Volume Segmentation for Radiotherapy Planning of Nasopharyngeal Carcinoma

Xiangde Luo, Jia Fu, Yunxin Zhong +39

Radiation therapy is a primary and effective NasoPharyngeal Carcinoma (NPC) treatment strategy. The precise delineation of Gross Tumor Volumes (GTVs) and Organs-At-Risk (OARs) is c…

eess.IV2025

MSWAL: 3D Multi-class Segmentation of Whole Abdominal Lesions Dataset

Zhaodong Wu, Qiaochu Zhao, Ming Hu +13

With the significantly increasing incidence and prevalence of abdominal diseases, there is a need to embrace greater use of new innovations and technology for the diagnosis and tre…

cs.CV2026

MMRareBench: A Rare-Disease Multimodal and Multi-Image Medical Benchmark

Junzhi Ning, Jiashi Lin, Yingying Fang +9

Multimodal large language models (MLLMs) have advanced clinical tasks for common conditions, but their performance on rare diseases remains largely untested. In rare-disease scenar…

cs.CV2025

S2-UniSeg: Fast Universal Agglomerative Pooling for Scalable Segment Anything without Supervision

Huihui Xu, Jin Ye, Hongqiu Wang +10

Recent self-supervised image segmentation models have achieved promising performance on semantic segmentation and class-agnostic instance segmentation. However, their pretraining s…

cs.CV2025

Scaling Tumor Segmentation: Best Lessons from Real and Synthetic Data

Qi Chen, Xinze Zhou, Chen Liu +11

AI for tumor segmentation is limited by the lack of large, voxel-wise annotated datasets, which are hard to create and require medical experts. In our proprietary JHH dataset of 3,…

eess.IV2023

Artifact Restoration in Histology Images with Diffusion Probabilistic Models

Zhenqi He, Junjun He, Jin Ye +1

Histological whole slide images (WSIs) can be usually compromised by artifacts, such as tissue folding and bubbles, which will increase the examination difficulty for both patholog…

eess.IV2026

Unified Medical Image Tokenizer for Autoregressive Synthesis and Understanding

Chenglong Ma, Yuanfeng Ji, Jin Ye +9

Autoregressive modeling has driven major advances in multimodal AI, yet its application to medical imaging remains constrained by the absence of a unified image tokenizer that simu…

cs.CV2023

Pick the Best Pre-trained Model: Towards Transferability Estimation for Medical Image Segmentation

Yuncheng Yang, Meng Wei, Junjun He +3

Transfer learning is a critical technique in training deep neural networks for the challenging medical image segmentation task that requires enormous resources. With the abundance…

cs.CV2026

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models

Yi Xin, Siqi Luo, Tianxiang Xu +13

Diffusion Multi-modal Large Language Models (dMLLMs) have recently emerged as a novel architecture unifying image generation and understanding. However, developing effective and ef…

cs.CV2024

PitVis-2023 Challenge: Workflow Recognition in videos of Endoscopic Pituitary Surgery

Adrito Das, Danyal Z. Khan, Dimitrios Psychogyios +29

The field of computer vision applied to videos of minimally invasive surgery is ever-growing. Workflow recognition pertains to the automated recognition of various aspects of a sur…

cs.CV2025

Towards Dynamic 3D Reconstruction of Hand-Instrument Interaction in Ophthalmic Surgery

Ming Hu, Zhengdi Yu, Feilong Tang +7

Accurate 3D reconstruction of hands and instruments is critical for vision-based analysis of ophthalmic microsurgery, yet progress has been hampered by the lack of realistic, large…

cs.CV2023

Generative Model Based Noise Robust Training for Unsupervised Domain Adaptation

Zhongying Deng, Da Li, Junjun He +2

Target domain pseudo-labelling has shown effectiveness in unsupervised domain adaptation (UDA). However, pseudo-labels of unlabeled target domain data are inevitably noisy due to t…

cs.CV2025

MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Fanqing Meng, Lingxiao Du, Zongkai Liu +12

DeepSeek R1, and o1 have demonstrated powerful reasoning capabilities in the text domain through stable large-scale reinforcement learning. To enable broader applications, some wor…

cs.CV2023

Learning with Explicit Shape Priors for Medical Image Segmentation

Xin You, Junjun He, Jie Yang +1

Medical image segmentation is a fundamental task for medical image analysis and surgical planning. In recent years, UNet-based networks have prevailed in the field of medical image…

cs.CV2024

Interactive Medical Image Segmentation: A Benchmark Dataset and Baseline

Junlong Cheng, Bin Fu, Jin Ye +10

Interactive Medical Image Segmentation (IMIS) has long been constrained by the limited availability of large-scale, diverse, and densely annotated datasets, which hinders model gen…

physics.med-ph2022

OpenKBP-Opt: An international and reproducible evaluation of 76 knowledge-based planning pipelines

Aaron Babier, Rafid Mahmood, Binghao Zhang +56

We establish an open framework for developing plan optimization models for knowledge-based planning (KBP) in radiotherapy. Our framework includes reference plans for 100 patients w…

eess.IV2023

A-Eval: A Benchmark for Cross-Dataset Evaluation of Abdominal Multi-Organ Segmentation

Ziyan Huang, Zhongying Deng, Jin Ye +11

Although deep learning have revolutionized abdominal multi-organ segmentation, models often struggle with generalization due to training on small, specific datasets. With the recen…

cs.CV2023

Token Sparsification for Faster Medical Image Segmentation

Lei Zhou, Huidong Liu, Joseph Bae +3

Can we use sparse tokens for dense prediction, e.g., segmentation? Although token sparsification has been applied to Vision Transformers (ViT) to accelerate classification, it is s…

cs.CL2025

Multimodal Human-AI Synergy for Medical Imaging Quality Control: A Hybrid Intelligence Framework with Adaptive Dataset Curation and Closed-Loop Evaluation

Zhi Qin, Qianhui Gui, Mouxiao Bian +14

Medical imaging quality control (QC) is essential for accurate diagnosis, yet traditional QC methods remain labor-intensive and subjective. To address this challenge, in this study…

cs.CV2024

SAM-Med3D-MoE: Towards a Non-Forgetting Segment Anything Model via Mixture of Experts for 3D Medical Image Segmentation

Guoan Wang, Jin Ye, Junlong Cheng +5

Volumetric medical image segmentation is pivotal in enhancing disease diagnosis, treatment planning, and advancing medical research. While existing volumetric foundation models for…

cs.CV2026

MedProbeBench: Systematic Benchmarking at Deep Evidence Integration for Expert-level Medical Guideline

Jiyao Liu, Jianghan Shen, Sida Song +19

Recent advances in deep research systems enable large language models to retrieve, synthesize, and reason over large-scale external knowledge. In medicine, developing clinical guid…

cs.CV2025

Brain Foundation Models with Hypergraph Dynamic Adapter for Brain Disease Analysis

Zhongying Deng, Haoyu Wang, Ziyan Huang +6

Brain diseases, such as Alzheimer's disease and brain tumors, present profound challenges due to their complexity and societal impact. Recent advancements in brain foundation model…

cs.CV2025

F^2TTA: Free-Form Test-Time Adaptation on Cross-Domain Medical Image Classification via Image-Level Disentangled Prompt Tuning

Wei Li, Jingyang Zhang, Lihao Liu +4

Test-Time Adaptation (TTA) has emerged as a promising solution for adapting a source model to unseen medical sites using unlabeled test data, due to the high cost of data annotatio…

quant-ph2024

Tunable ultrahigh reflection with broadband via collective atom-atom interaction in waveguide-QED system

Xin Wang, Junjun He, Zeyang Liao +1

We present a scheme for achieving broadband complete reflection by constructing photonic bandgap via collective atom-atom interaction in a one-dimensional (1D) waveguide quantum el…

cs.CV2024

SAM-Med3D: Towards General-purpose Segmentation Models for Volumetric Medical Images

Haoyu Wang, Sizheng Guo, Jin Ye +11

Existing volumetric medical image segmentation models are typically task-specific, excelling at specific target but struggling to generalize across anatomical structures or modalit…

cs.CV2024

OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Qingyun Li, Zhe Chen, Weiyun Wang +37

Image-text interleaved data, consisting of multiple images and texts arranged in a natural document format, aligns with the presentation paradigm of internet data and closely resem…

eess.IV2026

Open World MRI Reconstruction with Bias-Calibrated Adaptation

Jiyao Liu, Shangqi Gao, Lihao Liu +5

Real-world MRI reconstruction systems face the open-world challenge: test data from unseen imaging centers, anatomical structures, or acquisition protocols can differ drastically f…

cs.AI2025

HealthFlow: A Self-Evolving AI Agent with Meta Planning for Autonomous Healthcare Research

Yinghao Zhu, Yifan Qi, Zixiang Wang +8

The rapid proliferation of scientific knowledge presents a grand challenge: transforming this vast repository of information into an active engine for discovery, especially in high…

cs.CV2025

AOR: Anatomical Ontology-Guided Reasoning for Medical Large Multimodal Model in Chest X-Ray Interpretation

Qingqiu Li, Zihang Cui, Seongsu Bae +8

Chest X-rays (CXRs) are the most frequently performed imaging examinations in clinical settings. Recent advancements in Large Multimodal Models (LMMs) have enabled automated CXR in…

q-bio.BM2024

TourSynbio: A Multi-Modal Large Model and Agent Framework to Bridge Text and Protein Sequences for Protein Engineering

Yiqing Shen, Zan Chen, Michail Mamalakis +6

The structural similarities between protein sequences and natural languages have led to parallel advancements in deep learning across both domains. While large language models (LLM…

cs.CV2022

Dynamic Instance Domain Adaptation

Zhongying Deng, Kaiyang Zhou, Da Li +3

Most existing studies on unsupervised domain adaptation (UDA) assume that each domain's training samples come with domain labels (e.g., painting, photo). Samples from each domain a…

cs.CV2025

Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding

Yi Xin, Qi Qin, Siqi Luo +29

We introduce Lumina-DiMOO, an open-source foundational model for seamless multi-modal generation and understanding. Lumina-DiMOO sets itself apart from prior unified models by util…

eess.IV2025

Ophora: A Large-Scale Data-Driven Text-Guided Ophthalmic Surgical Video Generation Model

Wei Li, Ming Hu, Guoan Wang +7

In ophthalmic surgery, developing an AI system capable of interpreting surgical videos and predicting subsequent operations requires numerous ophthalmic surgical videos with high-q…

cs.CL2023

Enhancing Medical Task Performance in GPT-4V: A Comprehensive Study on Prompt Engineering Strategies

Pengcheng Chen, Ziyan Huang, Zhongying Deng +6

OpenAI's latest large vision-language model (LVLM), GPT-4V(ision), has piqued considerable interest for its potential in medical applications. Despite its promise, recent studies a…

cs.CV2026

UniMedVL: Unifying Medical Multimodal Understanding and Generation through Observation-Knowledge-Analysis

Junzhi Ning, Wei Li, Cheng Tang +24

Medical workflows routinely combine reading images with producing visual and textual outputs, making both image understanding and generation central to medical AI. Most existing sy…

cs.CV2026

StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding

Haolin Yang, Feilong Tang, Lingxiao Zhao +10

Real-time streaming video understanding in domains such as autonomous driving and intelligent surveillance poses challenges beyond conventional offline video processing, requiring…

cs.CV2023

SAM-Med2D

Junlong Cheng, Jin Ye, Zhongying Deng +12

The Segment Anything Model (SAM) represents a state-of-the-art research advancement in natural image segmentation, achieving impressive results with input prompts such as points an…

cs.CL2025

Silence is Not Consensus: Disrupting Agreement Bias in Multi-Agent LLMs via Catfish Agent for Clinical Decision Making

Yihan Wang, Qiao Yan, Zhenghao Xing +5

Large language models (LLMs) have demonstrated strong potential in clinical question answering, with recent multi-agent frameworks further improving diagnostic accuracy via collabo…

cs.AI2026

CuSearch: Curriculum Rollout Sampling via Search Depth for Agentic RAG

Jianghan Shen, Siqi Luo, Xinyu Cheng +6

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for training agentic retrieval-augmented generation (RAG) systems from outcome-only superv…

quant-ph2026

Theory of the Collective Many-body Subradiance in Waveguide Quantum Electrodynamics

Xin Wang, Junjun He, Zeyang Liao

We present an analytical theory for the most subradiant modes in a finite one-dimensional emitter array coupled to either an ideal or a nonideal waveguide. Using an effective non-H…

cs.CV2025

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Weiyun Wang, Zhangwei Gao, Lixin Gu +72

We introduce InternVL 3.5, a new family of open-source multimodal models that significantly advances versatility, reasoning capability, and inference efficiency along the InternVL…

cs.LG2025

Intern-S1: A Scientific Multimodal Foundation Model

Lei Bai, Zhongrui Cai, Yuhang Cao +173

In recent years, a plethora of open-source foundation models have emerged, achieving remarkable progress in some widely attended fields, with performance being quite close to that…

eess.IV2025

Towards Interpretable Counterfactual Generation via Multimodal Autoregression

Chenglong Ma, Yuanfeng Ji, Jin Ye +6

Counterfactual medical image generation enables clinicians to explore clinical hypotheses, such as predicting disease progression, facilitating their decision-making. While existin…

cs.LG2025

From AI for Science to Agentic Science: A Survey on Autonomous Scientific Discovery

Jiaqi Wei, Yuejin Yang, Xiang Zhang +24

Artificial intelligence (AI) is reshaping scientific discovery, evolving from specialized computational tools into autonomous research partners. We position Agentic Science as a pi…

eess.IV2025

Multimodal Medical Image Binding via Shared Text Embeddings

Yunhao Liu, Suyang Xi, Shiqi Liu +6

Medical image analysis increasingly relies on the integration of multiple imaging modalities to capture complementary anatomical and functional information, enabling more accurate…

cs.CV2026

MedQ-UNI: Toward Unified Medical Image Quality Assessment and Restoration via Vision-Language Modeling

Jiyao Liu, Junzhi Ning, Wanying Qu +4

Existing medical image restoration (Med-IR) methods are typically modality-specific or degradation-specific, failing to generalize across the heterogeneous degradations encountered…

cs.CV2024

OpenMEDLab: An Open-source Platform for Multi-modality Foundation Models in Medicine

Xiaosong Wang, Xiaofan Zhang, Guotai Wang +17

The emerging trend of advancing generalist artificial intelligence, such as GPTv4 and Gemini, has reshaped the landscape of research (academia and industry) in machine learning and…

cs.CV2024

OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining

Ming Hu, Kun Yuan, Yaling Shen +17

Surgical practice involves complex visual interpretation, procedural skills, and advanced medical knowledge, making surgical vision-language pretraining (VLP) particularly challeng…

eess.IV2024

Prompted Contextual Transformer for Incomplete-View CT Reconstruction

Chenglong Ma, Zilong Li, Junjun He +3

Incomplete-view computed tomography (CT) can shorten the data acquisition time and allow scanning of large objects, including sparse-view and limited-angle scenarios, each with var…

cs.AI2025

THE-Tree: Can Tracing Historical Evolution Enhance Scientific Verification and Reasoning?

Xin Wang, Jiyao Liu, Yulong Xiao +5

Large Language Models (LLMs) are accelerating scientific idea generation, but rigorously evaluating these numerous, often superficial, AI-generated propositions for novelty and fac…

cs.CV2026

HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding

Wenhui Liao, Hongliang Li, Pengyu Xie +15

Document parsing is a fundamental task in multimodal understanding, supporting a wide range of downstream applications such as information extraction and intelligent document analy…

eess.IV2024

SegBook: A Simple Baseline and Cookbook for Volumetric Medical Image Segmentation

Jin Ye, Ying Chen, Yanjun Li +7

Computed Tomography (CT) is one of the most popular modalities for medical imaging. By far, CT images have contributed to the largest publicly available datasets for volumetric med…

cs.CV2020

EfficientFCN: Holistically-guided Decoding for Semantic Segmentation

Jianbo Liu, Junjun He, Jiawei Zhang +2

Both performance and efficiency are important to semantic segmentation. State-of-the-art semantic segmentation algorithms are mostly based on dilated Fully Convolutional Networks (…

cs.CV2025

GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI

Tianbin Li, Yanzhou Su, Wei Li +15

Despite significant advancements in general AI, its effectiveness in the medical domain is limited by the lack of specialized medical knowledge. To address this, we formulate GMAI-…

cs.CV2025

HyperST: Hierarchical Hyperbolic Learning for Spatial Transcriptomics Prediction

Chen Zhang, Yilu An, Ying Chen +7

Spatial Transcriptomics (ST) merges the benefits of pathology images and gene expression, linking molecular profiles with tissue structure to analyze spot-level function comprehens…

cs.CL2026

Towards Efficient Medical Reasoning with Minimal Fine-Tuning Data

Xinlin Zhuang, Feilong Tang, Haolin Yang +9

Supervised Fine-Tuning (SFT) of the language backbone plays a pivotal role in adapting Vision-Language Models (VLMs) to specialized domains such as medical reasoning. However, exis…

cs.CV2025

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Jinguo Zhu, Weiyun Wang, Zhe Chen +48

We introduce InternVL3, a significant advancement in the InternVL series featuring a native multimodal pre-training paradigm. Rather than adapting a text-only large language model…

cs.CL2025

MMRC: A Large-Scale Benchmark for Understanding Multimodal Large Language Model in Real-World Conversation

Haochen Xue, Feilong Tang, Ming Hu +13

Recent multimodal large language models (MLLMs) have demonstrated significant potential in open-ended conversation, generating more accurate and personalized responses. However, th…

cs.CL2024

A Survey for Large Language Models in Biomedicine

Chong Wang, Mengyao Li, Junjun He +14

Recent breakthroughs in large language models (LLMs) offer unprecedented natural language understanding and generation capabilities. However, existing surveys on LLMs in biomedicin…

eess.IV2021

Group Shift Pointwise Convolution for Volumetric Medical Image Segmentation

Junjun He, Jin Ye, Cheng Li +5

Recent studies have witnessed the effectiveness of 3D convolutions on segmenting volumetric medical images. Compared with the 2D counterparts, 3D convolutions can capture the spati…

cs.CL2025

Benchmarking Chinese Medical LLMs: A Medbench-based Analysis of Performance Gaps and Hierarchical Optimization Strategies

Luyi Jiang, Jiayuan Chen, Lu Lu +4

The evaluation and improvement of medical large language models (LLMs) are critical for their real-world deployment, particularly in ensuring accuracy, safety, and ethical alignmen…

cs.CV2018

Prostate Segmentation using 2D Bridged U-net

Wanli Chen, Yue Zhang, Junjun He +4

In this paper, we focus on three problems in deep learning based medical image segmentation. Firstly, U-net, as a popular model for medical image segmentation, is difficult to trai…

cs.CV2025

HyperPath: Knowledge-Guided Hyperbolic Semantic Hierarchy Modeling for WSI Analysis

Peixiang Huang, Yanyan Huang, Weiqin Zhao +2

Pathology is essential for cancer diagnosis, with multiple instance learning (MIL) widely used for whole slide image (WSI) analysis. WSIs exhibit a natural hierarchy -- patches, re…

eess.IV2025

Multi-modal MRI Translation via Evidential Regression and Distribution Calibration

Jiyao Liu, Shangqi Gao, Yuxin Li +9

Multi-modal Magnetic Resonance Imaging (MRI) translation leverages information from source MRI sequences to generate target modalities, enabling comprehensive diagnosis while overc…

q-bio.QM2024

A Fine-tuning Dataset and Benchmark for Large Language Models for Protein Understanding

Yiqing Shen, Zan Chen, Michail Mamalakis +6

The parallels between protein sequences and natural language in their sequential structures have inspired the application of large language models (LLMs) to protein understanding.…

cs.CV2025

Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?

Pedro R. A. S. Bassi, Wenxuan Li, Yucheng Tang +50

How can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified…

cs.AI2026

Evidence-Grounded AI for Musculoskeletal Care

Wenjie Li, Yujie Zhang, Fanrui Zhang +34

The paper presents OrthoPilot, a clinical AI system powered by a large language model that integrates real-time hospital data and external medical knowledge to provide evidence‑bas…

#musculoskeletal care#clinical decision support#large language models#longitudinal patient management
cs.CV2025

FCN+: Global Receptive Convolution Makes FCN Great Again

Xiaoyu Ren, Zhongying Deng, Jin Ye +2

Fully convolutional network (FCN) is a seminal work for semantic segmentation. However, due to its limited receptive field, FCN cannot effectively capture global context informatio…

cs.CV2023

Vision Transformer Adapter for Dense Predictions

Zhe Chen, Yuchen Duan, Wenhai Wang +4

This work investigates a simple yet powerful dense prediction task adapter for Vision Transformer (ViT). Unlike recently advanced variants that incorporate vision-specific inductiv…

cs.CV2026

MedQ-Engine: A Closed-Loop Data Engine for Evolving MLLMs in Medical Image Quality Assessment

Jiyao Liu, Junzhi Ning, Wanying Qu +4

Medical image quality assessment (Med-IQA) is a prerequisite for clinical AI deployment, yet multimodal large language models (MLLMs) still fall substantially short of human expert…

cs.CL2025

A Survey of Scientific Large Language Models: From Data Foundations to Agent Frontiers

Ming Hu, Chenglong Ma, Wei Li +117

Scientific Large Language Models (Sci-LLMs) are transforming how knowledge is represented, integrated, and applied in scientific research, yet their progress is shaped by the compl…

cs.CV2026

MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models

Ryan D'Cunha, Alejandro Lozano, Xiaoxiao Sun +17

Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, from detecting lesions in chest X-rays to profiling cellular features in microscop…

cs.CV2026

TopoSculpt: Betti-Steered Topological Sculpting of 3D Fine-grained Tubular Shapes

Minghui Zhang, Yaoyu Liu, Junyang Wu +4

Medical tubular anatomical structures are inherently three-dimensional conduits with lumens, enclosing walls, and complex branching topologies. Accurate reconstruction of their geo…

cs.CV2025

SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image Understanding

Ying Chen, Guoan Wang, Yuanfeng Ji +7

Despite the progress made by multimodal large language models (MLLMs) in computational pathology, they remain limited by a predominant focus on patch-level analysis, missing essent…

cs.CL2025

TCM-3CEval: A Triaxial Benchmark for Assessing Responses from Large Language Models in Traditional Chinese Medicine

Tianai Huang, Lu Lu, Jiayuan Chen +5

Large language models (LLMs) excel in various NLP tasks and modern medicine, but their evaluation in traditional Chinese medicine (TCM) is underexplored. To address this, we introd…

cs.CV2024

Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

Zhangwei Gao, Zhe Chen, Erfei Cui +12

Multimodal large language models (MLLMs) have demonstrated impressive performance in vision-language tasks across a broad spectrum of domains. However, the large model scale and as…

cs.CV2026

Project Imaging-X: A Survey of 1000+ Open-Access Medical Imaging Datasets for Foundation Model Development

Zhongying Deng, Cheng Tang, Ziyan Huang +124

Foundation models have demonstrated remarkable success across diverse domains and tasks, primarily due to the thrive of large-scale, diverse, and high-quality datasets. However, in…

cs.CV2026

EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic Retrieval

Jiashi Lin, Changhong Jiang, Xiangru Lin +12

The paper proposes EvoGraph-R1, a framework that lets a retrieval agent dynamically evolve multimodal knowledge hypergraphs through actions like retrieval, web search, and graph ed…

#retrieval-augmented generation#dynamic knowledge graphs#graph-based reasoning#multimodal VQA
cs.LG2020

MIA-Prognosis: A Deep Learning Framework to Predict Therapy Response

Jiancheng Yang, Jiajun Chen, Kaiming Kuang +3

Predicting clinical outcome is remarkably important but challenging. Research efforts have been paid on seeking significant biomarkers associated with the therapy response or/and p…

cs.CV2026

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning

Cheng Tang, Junzhi Ning, Min Cen +9

The paper presents SIVA-RL, a framework that uses sample-wise visual interventions to align sensitivity and invariance in multimodal reinforcement learning models, leading to bette…

#multimodal reinforcement learning#visual grounding#policy alignment#visual intervention
cs.CV2025

SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Dongyang Liu, Renrui Zhang, Longtian Qiu +16

We propose SPHINX-X, an extensive Multimodality Large Language Model (MLLM) series developed upon SPHINX. To improve the architecture and training efficiency, we modify the SPHINX…

eess.IV2025

StructDiff: Structure-aware Diffusion Model for 3D Fine-grained Medical Image Synthesis

Jiahao Xia, Yutao Hu, Yaolei Qi +6

Solving medical imaging data scarcity through semantic image generation has attracted growing attention in recent years. However, existing generative models mainly focus on synthes…

cs.CV2026

TopoVST: Toward Topology-fidelitous Vessel Skeleton Tracking

Yaoyu Liu, Minghui Zhang, Junjun He +1

Automatic extraction of vessel skeletons is crucial for many clinical applications. However, achieving topologically faithful delineation of thin vessel skeletons remains highly ch…

eess.IV2025

ADAgent: LLM Agent for Alzheimer's Disease Analysis with Collaborative Coordinator

Wenlong Hou, Guangqian Yang, Ye Du +5

Alzheimer's disease (AD) is a progressive and irreversible neurodegenerative disease. Early and precise diagnosis of AD is crucial for timely intervention and treatment planning to…

cs.CL2025

A Novel Ophthalmic Benchmark for Evaluating Multimodal Large Language Models with Fundus Photographs and OCT Images

Xiaoyi Liang, Mouxiao Bian, Moxin Chen +4

In recent years, large language models (LLMs) have demonstrated remarkable potential across various medical applications. Building on this foundation, multimodal large language mod…

cs.CV2020

Tensor Low-Rank Reconstruction for Semantic Segmentation

Wanli Chen, Xinge Zhu, Ruoqi Sun +4

Context information plays an indispensable role in the success of semantic segmentation. Recently, non-local self-attention based methods are proved to be effective for context inf…

cs.CV2023

STU-Net: Scalable and Transferable Medical Image Segmentation Models Empowered by Large-Scale Supervised Pre-training

Ziyan Huang, Haoyu Wang, Zhongying Deng +8

Large-scale models pre-trained on large-scale datasets have profoundly advanced the development of deep learning. However, the state-of-the-art models for medical image segmentatio…

cs.CV2025

EventRR: Event Referential Reasoning for Referring Video Object Segmentation

Huihui Xu, Jiashi Lin, Haoyu Chen +2

Referring Video Object Segmentation (RVOS) aims to segment out the object in a video referred by an expression. Current RVOS methods view referring expressions as unstructured sequ…

cs.AI2025

Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows

Wanghan Xu, Yuhao Zhou, Yifan Zhou +104

Despite advances in scientific AI, a coherent framework for Scientific General Intelligence (SGI)-the ability to autonomously conceive, investigate, and reason across scientific do…