From the 1 of 2 linked papers with an AI index.
2 papers
cs.CL2026
Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces
Manas Pathak, Xingyao Chen, Shuozhe Li +2
The paper introduces the Filtered Reasoning Score (FRS), a metric that evaluates the quality of reasoning traces from large language models by focusing on the most confident genera…
cs.CL2025
Beyond Query-Level Comparison: Fine-Grained Reinforcement Learning for Text-to-SQL with Automated Interpretable Critiques
Guifeng Wang, Yuanfeng Song, Meng Yang +3
Text-to-SQL, a pivotal natural language processing (NLP) task that converts textual queries into executable SQL, has seen substantial progress in recent years. However, existing ev…