activity
20242026
collaborators

8 papers

stat.ME2026

Reusing Operational Evidence After Context Changes: A Conservative Bayesian Framework for Autonomous Vehicle Safety

Robab Aghazadeh Chakherlou, Siddartha Khastgir, Xingyu Zhao

Operational evidence, i.e., evidence of operation without failure is an important component of confidence in the safety or reliability of a system in service, but it is costly to c…

stat.ML2026

Modeling Memory-Dependent Reliability of LLMs: A Hidden Markov Model

Robab Aghazadeh Chakherlou, Siddartha Khastgir, Peter Popov +1

Reliability assessment of large language models (LLMs) seeks to estimate the probability that a model produces correct responses under a specified operational profile. Conventional…

cs.CV2026

Probabilistic Robustness in Medical Image Classification

Yi Zhang, Siddartha Khastgir, Xingyu Zhao

Deep learning (DL) has shown strong performance in medical image classification, but its trustworthy deployment remains challenging in safety-critical clinical settings, where pred…

cs.AI2026

Quantifying Fidelity: A Decisive Feature Approach to Comparing Synthetic and Real Imagery

Danial Safaei, Siddartha Khastgir, Mohsen Alirezaei +4

Virtual testing using synthetic data has become a cornerstone of autonomous vehicle (AV) safety assurance. Despite progress in improving visual realism through advanced simulators…

cs.AI2025

Uncertainty-Aware Measurement of Scenario Suite Representativeness for Autonomous Systems

Robab Aghazadeh Chakherlou, Siddartha Khastgir, Xingyu Zhao +2

Assuring the trustworthiness and safety of AI systems, e.g., autonomous vehicles (AV), depends critically on the data-related safety properties, e.g., representativeness, completen…

cs.SE2025

A Hierarchical Imprecise Probability Approach to Reliability Assessment of Large Language Models

Robab Aghazadeh-Chakherlou, Qing Guo, Siddartha Khastgir +3

Large Language Models (LLMs) are increasingly deployed across diverse domains, raising the need for rigorous reliability assessment methods. Existing benchmark-based evaluations pr…