1 paper · 1 filter
Xing Zi, Xinying Zhou, Jinghao Xiao +2
While Large Language Models (LLMs) achieve expert-level performance on standard medical benchmarks through single-hop factual recall, they severely struggle with the complex, multi…