168 citations
- Icahn School of Medicine at Mount SinaiUS4 papers
- Indiana University – Purdue University IndianapolisUS4 papers
- Institute for Research in Fundamental SciencesIR4 papers
- Cedars-Sinai Medical CenterUS3 papers
- Child Health and Development InstituteUS3 papers
- Research Institute for Endocrine SciencesIR3 papers
- Sharif University of TechnologyIR3 papers
- Arizona State UniversityUS2 papers
- Columbia UniversityUS2 papers
- Ontario Tech UniversityCA2 papers
- Shahid Beheshti UniversityIR2 papers
- Shiraz University of Medical SciencesIR2 papers
20 papers
Safety boundary maintenance in consumer AI systems responding to pediatric health queries: a cross-platform benchmark evaluation under naturalistic and adversarially pressured conditions
Vahideh Zolfaghari, Leila Mashhadi, Mitra Ahadi +2
Consumer artificial intelligence chatbots are now accessed by hundreds of millions of users seeking health information, yet systematic evaluation of their safety boundary maintenan…
Ultrasound-based detection and malignancy prediction of breast lesions eligible for biopsy: A multi-center clinical-scenario study using nomograms, large language models, and radiologist evaluation
Ali Abbasian Ardakani, Afshin Mohammadi, Taha Yusuf Kuzan +7
To develop and externally validate integrated ultrasound nomograms combining BIRADS features and quantitative morphometric characteristics, and to compare their performance with ex…
Vision Language Models versus Machine Learning Models Performance on Polyp Detection and Classification in Colonoscopy Images
Mohammad Amin Khalafi, Seyed Amir Ahmad Safavi-Naini, Ameneh Salehi +13
Introduction: This study provides a comprehensive performance assessment of vision-language models (VLMs) against established convolutional neural networks (CNNs) and classic machi…
Self-Reported Confidence of Large Language Models in Gastroenterology: Analysis of Commercial, Open-Source, and Quantized Models
Nariman Naderi, Seyed Amir Ahmad Safavi-Naini, Thomas Savage +4
This study evaluated self-reported response certainty across several large language models (GPT, Claude, Llama, Phi, Mistral, Gemini, Gemma, and Qwen) using 300 gastroenterology bo…
Vision-Language and Large Language Model Performance in Gastroenterology: GPT, Claude, Llama, Phi, Mistral, Gemma, and Quantized Models
Seyed Amir Ahmad Safavi-Naini, Shuhaib Ali, Omer Shahab +15
Background and Aims: This study evaluates the medical reasoning performance of large language models (LLMs) and vision language models (VLMs) in gastroenterology. Methods: We used…
Large Language Models versus Classical Machine Learning: Performance in COVID-19 Mortality Prediction Using High-Dimensional Tabular Data
Mohammadreza Ghaffarzadeh-Esfahani, Mahdi Ghaffarzadeh-Esfahani, Arian Salahi-Niri +39
This study compared the performance of classical feature-based machine learning models (CMLs) and large language models (LLMs) in predicting COVID-19 mortality using high-dimension…