2 papers
cs.CL2026
Performance of Clinical AI System and Physicians and Frontier Language Models in primary care diagnostics
Andy Nkansah, Hanna Plotnitskaya, Stanislau Salavei +6
Clinical AI evaluation should encompass diagnosis and management after adaptive information gathering. We compared Doctorina, eight physicians and four standalone frontier language…
cs.CL2026
Doctorina MedBench: A Dialogue-Based Benchmark and Evaluation Framework for Agent-Based Medical AI
Anna Kozlova, Stanislau Salavei, Pavel Satalkin +3
We present Doctorina MedBench, an evaluation framework for agent-based medical AI based on the simulation of physician-patient interactions. Unlike traditional medical benchmarks t…