3 papers
cs.CR2026
Full-range Binary Classifier Calibration for Stable Model Updates in Production
Konstantin Berlin
Detection models running in adversarial environments face a malicious distribution that drifts rapidly while the benign distribution stays comparatively stable, so teams retrain an…
cs.CL2026
Improving Labeling Consistency with Detailed Constitutional Definitions and AI-Driven Evaluation
Konstantin Berlin, Adam Swanda
Many automated labeling pipelines classify inputs into categories defined by a written specification, content moderation being a prominent use case. Simple category definitions are…
cs.CR2025
A Framework for Rapidly Developing and Deploying Protection Against Large Language Model Attacks
Adam Swanda, Amy Chang, Alexander Chen +3
The widespread adoption of Large Language Models (LLMs) has revolutionized AI deployment, enabling autonomous and semi-autonomous applications across industries through intuitive l…