4 papers
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…
Fair-Eye Net: A Fair, Trustworthy, Multimodal Integrated Glaucoma Full Chain AI System
Wenbin Wei, Suyuan Yao, Cheng Huang +1
Glaucoma is a top cause of irreversible blindness globally, making early detection and longitudinal follow-up pivotal to preventing permanent vision loss. Current screening and pro…
Foundation Models for Clinical Records at Health System Scale
Haresh Rengaraj Rajamohan, Xiang Gao, Weicheng Zhu +5
Large-scale pretraining has transformed modeling of language and other data types, but its potential remains underexplored in healthcare with structured electronic health records (…
RefSAM3D: Adapting SAM with Cross-modal Reference for 3D Medical Image Segmentation
Xiang Gao, Kai Lu
The Segment Anything Model (SAM), originally built on a 2D Vision Transformer (ViT), excels at capturing global patterns in 2D natural images but struggles with 3D medical imaging…