2 papers
cs.CR2026
Optimus: A Robust Defense Framework for Mitigating Toxicity while Fine-Tuning Conversational AI
Aravind Cheruvu, Shravya Kanchi, Sifat Muhammad Abdullah +4
Customizing Large Language Models (LLMs) on untrusted datasets poses severe risks of injecting toxic behaviors. In this work, we introduce Optimus, a novel defense framework design…
cs.AI2026
Judge Reliability Harness: Stress Testing the Reliability of LLM Judges
Sunishchal Dev, Andrew Sloan, Joshua Kavner +2
We present the Judge Reliability Harness, an open source library for constructing validation suites that test the reliability of LLM judges. As LLM based scoring is widely deployed…