2 papers
cs.CV2026
PRBench: A Standardized Probabilistic Robustness Benchmark
Yi Zhang, Zheng Wang, Zhen Chen +5
Deep learning models are notoriously vulnerable to imperceptible perturbations. Most existing research centers on adversarial robustness (AR), which evaluates models under worst-ca…
cs.SE2026
A Hierarchical Imprecise Probability Approach to Reliability Assessment of Large Language Models
Robab Aghazadeh-Chakherlou, Qing Guo, Siddartha Khastgir +3
Large Language Models (LLMs) are increasingly deployed across diverse domains, raising the need for rigorous reliability assessment methods. Existing benchmark-based evaluations pr…