4 papers
GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration
Junjie Zhao, Jingyi Liang, Zhenyang Cai +22
While large language models (LLMs) hold transformative potential for medicine, their reasoning robustness and safety in real-world clinical scenarios remain critically underexplore…
MicroVerse: A Preliminary Exploration Toward a Micro-World Simulation
Rongsheng Wang, Minghao Wu, Hongru Zhou +4
Recent advances in video generation have opened new avenues for macroscopic simulation of complex dynamic systems, but their application to microscopic phenomena remains largely un…
LiveClin: A Live Clinical Benchmark without Leakage
Xidong Wang, Shuqi Guo, Yue Shen +6
The reliability of medical LLM evaluation is critically undermined by data contamination and knowledge obsolescence, leading to inflated scores on static benchmarks. To address the…
DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry
Zhenyang Cai, Jiaming Zhang, Junjie Zhao +21
Reliable interpretation of multimodal data in dentistry is essential for automated oral healthcare, yet current multimodal large language models (MLLMs) struggle to capture fine-gr…