6 papers
When Does Restricting a Coding Agent to execute_code Help? A Regime Agent-Design Ablation
Hong Yang, Qi Yu, Travis Desell
Modern coding agents expose multiple tool surfaces -- IDE primitives, bash, and Model Context Protocol (MCP) code-execution -- and the field has shipped three contradictory claims…
Distributed Quality-Diversity Search for Toxicity in Large Language Models
Onkar Shelar, Travis Desell
Large Language Models remain vulnerable to adversarial prompts that elicit harmful responses, and scaling red-teaming to cover a broad range of failure modes is constrained by the…
GSC-QEMit: A Telemetry-Driven Hierarchical Forecast-and-Bandit Framework for Adaptive Quantum Error Mitigation
Steven Szachara, Sheeraja Rajakrishnan, Dylan Jay Van Allen +3
Quantum error mitigation (QEM) is essential for extracting reliable results from near-term quantum devices, yet practical deployments must balance mitigation strength against runti…
Diversifying Toxicity Search in Large Language Models Through Speciation
Onkar Shelar, Travis Desell
Evolutionary prompt search is a practical black-box approach for red teaming large language models, however existing methods often collapse onto a small family of high-performing p…
Beyond the Class Subspace: Teacher-Guided Training for Reliable Out-of-Distribution Detection in Single-Domain Models
Hong Yang, Devroop Kar, Qi Yu +2
Out-of-distribution (OOD) detection methods perform well on multi-domain benchmarks, yet many practical systems are trained on single-domain data. We show that this regime induces…
ToxSearch: Evolving Prompts for Toxicity Search in Large Language Models
Onkar Shelar, Travis Desell
Large Language Models remain vulnerable to adversarial prompts that elicit toxic content even after safety alignment. We present ToxSearch, a black-box evolutionary framework that…