9 papers
CAPO: Critic-Guided Action-Aligned Policy Optimization for Advancing LLM Agent Capabilities
Daoyu Wang, Qingchuan Li, Mingyue Cheng +6
Reinforcement learning (RL) has become a key technique for improving the agentic capabilities of large language models (LLMs). Although critic-free methods such as GRPO are increas…
GeoDecider: An Evidence-Grounded Agent for Geological Interpretation via Deliberative Reasoning
Jiahao Wang, Mingyue Cheng, Yitong Zhou +6
Geological interpretation infers subsurface properties and structures from indirect geophysical observations. Well-log classification provides a measurable setting by assigning geo…
Survey of Computerized Adaptive Testing: A Machine Learning Perspective
Yan Zhuang, Qi Liu, Haoyang Bi +12
Computerized Adaptive Testing (CAT) offers an efficient and personalized method for assessing examinee proficiency by dynamically adjusting test questions based on individual perfo…
CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive Perspective
Jiayu Liu, Zhenya Huang, Wei Dai +7
Although large language models (LLMs) show promise in solving complex mathematical tasks, existing evaluation paradigms rely solely on a coarse measure of overall answer accuracy,…
From Objectives to Questions: A Planning-based Framework for Educational Mathematical Question Generation
Cheng Cheng, Zhenya Huang, Guanhao Zhao +5
Automatically generating high-quality mathematical problems that align with educational objectives is a crucial task in NLP-based educational technology. Traditional generation met…
am-ELO: A Stable Framework for Arena-based LLM Evaluation
Zirui Liu, Jiatong Li, Yan Zhuang +5
Arena-based evaluation is a fundamental yet significant evaluation paradigm for modern AI models, especially large language models (LLMs). Existing framework based on ELO rating sy…