2 papers
cs.AI2025
Multidimensional Rubric-oriented Reward Model Learning via Geometric Projection Reference Constraints
Yongnan Jin, Xurui Li, Feng Cao +2
The integration of large language models (LLMs) into medical practice offers transformative potential, yet their real-world clinical applicability remains constrained by critical a…
cs.CL2025
SUPERChem: A Multimodal Reasoning Benchmark in Chemistry
Zehua Zhao, Zhixian Huang, Junren Li +28
Current benchmarks for evaluating the chemical reasoning capabilities of Large Language Models (LLMs) are limited by oversimplified tasks, lack of process-level evaluation, and mis…