2 papers
cs.AI2026
Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance
Alexander Thomas, Hubert P. H. Shum, Darren Nellis +4
The transport of dangerous goods by sea is a high-consequence activity governed by the International Maritime Dangerous Goods (IMDG) Code, a complex regulatory framework where erro…
cs.CL2026
A Multi-Domain Red Teaming Framework for Safety, Robustness, and Fairness Evaluation of Medical Large Language Models
Andrei Marian Feier, Veysel Kocaman, Yigit Gul +6
Large language models (LLMs) are increasingly deployed across healthcare, yet existing benchmarks fail to capture model behavior under adversarial or ethically complex conditions c…