most citedLingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning

5 citations · 6 across the 9 of their papers we have counts for

collaborators

9 papers

cs.CV2026

On the Design Fundamentals of Pixel Text Representation Learning

Chaohao Yuan, Ruifeng Yuan, Zhuoxu Huang +4

Text-rich visual inputs require models that can read, retrieve, and compress language directly in pixel space, yet existing pixel-text encoders struggle with fixed resolution pretr…

cs.CV2026

RadSight: Towards Perceptually Reliable Multimodal Radiology Image Understanding

Jianqin Liu, Weiwei Cao, Wanxing Chang +7

Medical multimodal large language models (MLLMs) are increasingly expected to perform complex image understanding tasks, yet their reliability is often compromised by frequent erro…

cs.CE2026

AtomiMed: Hierarchical Atomic Fact-Checking for Universal Clinical-Aware Medical Report Evaluation

Yuan Wang, Wanxing Chang, Songtao Jiang +8

Traditional metrics for Medical Report Generation (MRG) predominantly rely on surface-level n-gram overlap, which fails to capture clinical factual accuracy and often overlooks cat…

cs.CV2026

Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed Tomography

Bowen Shi, Weiwei Cao, Ruifeng Yuan +5

Vision-language pre-training (VLP) holds great promise for general-purpose medical AI by leveraging radiology reports as rich textual supervision, yet existing methods struggle wit…

cs.CL2026

Understanding the Behaviors of Environment-aware Information Retrieval

Ruifeng Yuan, Chaohao Yuan, David Dai +4

Recent retrieval-augmented generation (RAG) approaches have demonstrated strong capability in handling complex queries, yet current research overlooks a critical challenge: differe…

cs.AI2026★ 1 cited

CT-FineBench: A Diagnostic Fidelity Benchmark for Fine-Grained Evaluation of CT Report Generation

Ruifeng Yuan, Wanxing Chang, Weiwei Cao +4

The evaluation of generated reports remains a critical challenge in Computed Tomography (CT) report generation, due to the large volume of text, the diversity and complexity of fin…