1 paper · 1 filter
Robab Aghazadeh-Chakherlou, Qing Guo, Siddartha Khastgir +3
Large Language Models (LLMs) are increasingly deployed across diverse domains, raising the need for rigorous reliability assessment methods. Existing benchmark-based evaluations pr…