Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
SMART: Self-Generating and Self-Validating Multi-Dimensional Assessment for LLMs' Mathematical Problem Solving
Yujie Hou, Mei Wang, Yaoyao Zhong +3
Large Language Models (LLMs) have achieved remarkable performance across a wide range of mathematical benchmarks. However, concerns remain as to whether these successes reflect gen…
cs.AI2026
VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
Can Li, Ying Liu, Ting Zhang +2
Large Multimodal Models have achieved remarkable progress in integrating vision and language, enabling strong performance across perception, reasoning, and domain-specific tasks. H…