5 citations · 5 across the 3 of their papers we have counts for
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
GeoDecider: An Evidence-Grounded Agent for Geological Interpretation via Deliberative Reasoning
Jiahao Wang, Mingyue Cheng, Yitong Zhou +6
Geological interpretation infers subsurface properties and structures from indirect geophysical observations. Well-log classification provides a measurable setting by assigning geo…
cs.AI2025
CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive Perspective
Jiayu Liu, Zhenya Huang, Wei Dai +7
Although large language models (LLMs) show promise in solving complex mathematical tasks, existing evaluation paradigms rely solely on a coarse measure of overall answer accuracy,…
cs.AI2025
am-ELO: A Stable Framework for Arena-based LLM Evaluation
Zirui Liu, Jiatong Li, Yan Zhuang +5
Arena-based evaluation is a fundamental yet significant evaluation paradigm for modern AI models, especially large language models (LLMs). Existing framework based on ELO rating sy…