2 papers
cs.CL2026
Beyond Semantic Accuracy: Consequence-Aware Evaluation for Safety-Critical Language Understanding
Yujing Chang, Thinh Pham, Van-Phat Thai +4
Can language models be trusted in safety- critical operations? In such settings, strong per- formance on semantic metrics does not guaran- tee operational reliability: a misread al…
cs.CL2026
Safety-Oriented Evaluation of Language Understanding Systems for Air Traffic Control
Yujing Chang, Yash Guleria, Duc-Thinh Pham +4
Air Traffic Control (ATC) is a safety-critical domain in which incorrect interpretation of instructions may lead to severe operational consequences. While large language models (LL…