2 papers
cs.LG2026
Beyond Red-Teaming: Formal Guarantees of LLM Guardrail Classifiers
Nikita Kezins, Urbas Ekka, Pascal Berrang +1
Guardrail Classifiers defend production language models against harmful behavior, but although results seem promising in testing, they provide no formal guarantees. Providing forma…
cs.LG2025
Building Machine Learning Challenges for Anomaly Detection in Science
Elizabeth G. Campolongo, Yuan-Tang Chou, Ekaterina Govorkova +148
Scientific discoveries are often made by finding a pattern or object that was not predicted by the known rules of science. Oftentimes, these anomalous events or objects that do not…