chain of thought 1counterfactual testing 1large language models 1mathematical reasoning 1reasoning evaluation 1reference-free metrics 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CR2026
A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models
Vivek Shukla, Varun Shukla, Atul +2
The paper proposes the Reasoning Answer Faithfulness Score (RAFS), a reference‑free metric that evaluates whether a large language model’s mathematical chain‑of‑thought trace is lo…
cs.LG2026
Picking the Right Specialist: Attentive Neural Process-based Selection of Task-Specialized Models as Tools for Agentic Healthcare Systems
Pramit Saha, Joshua Strong, Mohammad Alsharid +2
Task-specialized models form the backbone of agentic healthcare systems, enabling the agents to answer clinical queries across tasks such as disease diagnosis, localization, and re…