3 citations · 3 across the 1 of their papers we have counts for
1 paper
Zehua Zhao, Zhixian Huang, Junren Li +28
Current benchmarks for evaluating the chemical reasoning capabilities of Large Language Models (LLMs) are limited by oversimplified tasks, lack of process-level evaluation, and mis…