Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Hidden Reliability Risks in Large Language Models: Systematic Identification of Precision-Induced Output Disagreements
Yifei Wang, Tianlin Li, Xiaohan Zhang +4
Large language models (LLMs) are increasingly deployed under diverse numerical precision configurations, including standard floating-point formats (e.g., bfloat16 and float16) and…
cs.AI2025
Decictor: Towards Evaluating the Robustness of Decision-Making in Autonomous Driving Systems
Mingfei Cheng, Yuan Zhou, Xiaofei Xie +3
Autonomous Driving System (ADS) testing is crucial in ADS development, with the current primary focus being on safety. However, the evaluation of non-safety-critical performance, p…