chain of thought 1counterfactual testing 1large language models 1mathematical reasoning 1reasoning evaluation 1reference-free metrics 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CR2026
A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models
Vivek Shukla, Varun Shukla, Atul +2
The paper proposes the Reasoning Answer Faithfulness Score (RAFS), a reference‑free metric that evaluates whether a large language model’s mathematical chain‑of‑thought trace is lo…
cs.DC2026
Secure and Low-Latency IoT Analytics Using an Edge-Based Streaming Architecture
Atul, Varun Shukla, Vivek Shukla +1
The rapid growth of Internet of Things (IoT) devices has led to large-scale continuous data streams that require realtime processing. Traditional cloud-centric architectures fail t…