2 papers
cs.AI2026
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation
Sahil Pardasani, Madhusudan Singh
LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims that DeepSeek R1 outperformed O…
cs.AI2026
Bridging the Know-Act Gap via Task-Level Autoregressive Reasoning
Jihyun Janice Ahn, Ryo Kamoi, Berk Atil +34
LLMs often generate seemingly valid answers to flawed or ill-posed inputs. This is not due to missing knowledge: under discriminative prompting, the same models can mostly identify…