2 papers
cs.AI2026
SFBench: The SciFy Scientific Feasibility Benchmark
Cash Costello, James Mayfield, Elsbeth Turcan +7
We present SFBench, a benchmark dataset for evaluating systems that assess the feasibility of scientific claims. SFBench includes 197 claims in materials science, each annotated wi…
cs.CV2025
Causality-Driven Audits of Model Robustness
Nathan Drenkow, William Paul, Chris Ribaudo +1
Robustness audits of deep neural networks (DNN) provide a means to uncover model sensitivities to the challenging real-world imaging conditions that significantly degrade DNN perfo…