5 papers
SMART: Self-Generating and Self-Validating Multi-Dimensional Assessment for LLMs' Mathematical Problem Solving
Yujie Hou, Mei Wang, Yaoyao Zhong +3
Large Language Models (LLMs) have achieved remarkable performance across a wide range of mathematical benchmarks. However, concerns remain as to whether these successes reflect gen…
VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
Can Li, Ying Liu, Ting Zhang +2
Large Multimodal Models have achieved remarkable progress in integrating vision and language, enabling strong performance across perception, reasoning, and domain-specific tasks. H…
SketchJudge: A Diagnostic Benchmark for Grading Hand-drawn Diagrams with Multimodal Large Language Models
Yuhang Su, Mei Wang, Yaoyao Zhong +4
While Multimodal Large Language Models (MLLMs) have achieved remarkable progress in visual understanding, they often struggle when faced with the unstructured and ambiguous nature…
Discerning minds or generic tutors? Evaluating instructional guidance capabilities in Socratic LLMs
Ying Liu, Can Li, Ting Zhang +4
The conversational capabilities of large language models hold significant promise for enabling scalable and interactive tutoring. While prior research has primarily examined their…
From Answers to Questions: EQGBench for Evaluating LLMs' Educational Question Generation
Chengliang Zhou, Mei Wang, Ting Zhang +3
Large Language Models (LLMs) have demonstrated remarkable capabilities in mathematical problem-solving. However, the transition from providing answers to generating high-quality ed…