5 papers
GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration
Junjie Zhao, Jingyi Liang, Zhenyang Cai +22
While large language models (LLMs) hold transformative potential for medicine, their reasoning robustness and safety in real-world clinical scenarios remain critically underexplore…
CRPO: Character-centric Group Relative Policy Optimization for Role-aware Reasoning in Role-playing Agents
Yihong Tang, Kehai Chen, Liang Yue +2
Recent advancements in Reinforcement Learning (RL), particularly Group Relative Policy Optimization (GRPO), have significantly enhanced the reasoning capabilities of Large Language…
HiMed: Incentivizing Hindi Reasoning in Medical LLMs
Dingfeng Jiang, Han Yan, Chenze Ma +12
Medical large language models hold promise for reducing healthcare disparities, yet Hindi remains severely underrepresented. While medical LLMs excel in high-resource languages, th…
GeoSym127K: Scalable Symbolically-verifiable Synthesis for Multimodal Geometric Reasoning
Jinhao Jing, Zheng Ma, Jinwei Liang +9
Large Multimodal Models (LMMs) often struggle with geometric reasoning due to visual hallucinations and a lack of mathematically precise Chain-of-Thought (CoT) data. To address thi…
Character-R1: Enhancing Role-Aware Reasoning in Role-Playing Agents via RLVR
Yihong Tang, Kehai Chen, Xuefeng Bai +4
Current role-playing agents (RPAs) are typically constructed by imitating surface-level behaviors, but this approach lacks internal cognitive consistency, often causing out-of-char…