Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Benchmarking Multi-turn Medical Diagnosis: Hold, Lure, and Self-Correction
Jinrui Fang, Runhan Chen, Xu Yang +9
Large language models (LLMs) achieve high accuracy in medical diagnosis when all clinical information is provided in a single turn, yet how they behave under multi-turn evidence ac…
cs.CL2025
Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language Models
Kyle Cox, Jiawei Xu, Yikun Han +6
An interesting behavior in large language models (LLMs) is prompt sensitivity. When provided with different but semantically equivalent versions of the same prompt, models may prod…