Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
SUPERChem: A Multimodal Reasoning Benchmark in Chemistry
Zehua Zhao, Zhixian Huang, Junren Li +28
Current benchmarks for evaluating the chemical reasoning capabilities of Large Language Models (LLMs) are limited by oversimplified tasks, lack of process-level evaluation, and mis…
cs.CL2025
We Need Knowledge Distillation for Solving Math Word Problems
Zhenquan Shen, Xinguo Yu, Xiaotian Cheng +2
The enhancement of mathematical capabilities in large language models (LLMs) fosters new developments in mathematics education within primary and secondary schools, particularly as…