From the 1 of 4 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces
Congchao Wang, Diwakar Singh, Qiaozi Gao +3
Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective reasoning data curation, strong training signals during reinforcement lear…
cs.AI2026
A Threshold Exceedance Framework for CBRN Uplift Evaluation in Frontier Language Models
Rahul Gupta, Abhinav Mohanty, Payal Motwani +8
The paper introduces a Threshold Exceedance Criteria (TEC) framework to systematically evaluate whether frontier language models increase a non‑expert's ability to plan chemical, b…