Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
Nearchos Potamitis, Vansh Ramani, Har Ashish Arora +3
Benchmark scores for LLM reasoning systems are reported as single numbers, yet the same model, strategy, and task can produce meaningfully different answers and costs across repeat…
cs.AI2026
Fully Open Meditron: An Auditable Pipeline for Clinical LLMs
Xavier Theimer-Lienhard, Mushtaha El-Amin, Fay Elhassan +5
Clinical decision support systems (CDSS) require scrutable, auditable pipelines that enable rigorous, reproducible validation. Yet current LLM-based CDSS remain largely opaque. Mos…