activity
20242026
collaborators

8 papers

cs.CV2026

Mitigating Batch Effects in Histopathology via Language-Mediated Robust Embedding Generation

Yishu Zhang, Shushan Wu, Zhenzhong Zhang +8

Pathology foundation models (PFMs) have demonstrated strong potential across clinical and scientific applications, yet their performance is often hindered by batch effects, which a…

cs.LG2026

MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning

Peng Xia, Jinglu Wang, Yibo Peng +10

Medical Large Vision-Language Models (Med-LVLMs) have shown strong potential in multimodal diagnostic tasks. However, existing single-agent models struggle to generalize across div…

cs.LG2025

Foundation Model in Biomedicine

Xiangrui Liu, Yuanyuan Zhang, Qianyu Shang +14

Foundation models, first introduced in 2021, refer to large-scale pretrained models (e.g., large language models (LLMs) and vision-language models (VLMs)) that learn from extensive…

cs.CV2025

MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization

Kangyu Zhu, Peng Xia, Yun Li +3

The advancement of Large Vision-Language Models (LVLMs) has propelled their application in the medical field. However, Medical LVLMs (Med-LVLMs) encounter factuality challenges due…

q-bio.GN2025

CellTypeAgent: Trustworthy cell type annotation with Large Language Models

Jiawen Chen, Jianghao Zhang, Huaxiu Yao +1

Cell type annotation is a critical yet laborious step in single-cell RNA sequencing analysis. We present a trustworthy large language model (LLM)-agent, CellTypeAgent, which integr…

cs.LG2025

MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding

Siwei Han, Peng Xia, Ruiyi Zhang +4

Document Question Answering (DocQA) is a very common task. Existing methods using Large Language Models (LLMs) or Large Vision Language Models (LVLMs) and Retrieval Augmented Gener…