3 papers
cs.CY2026
The Doctor Will (Still) See You Now: On the Structural Limits of Agentic AI in Healthcare
Gabriela Aránguiz Dias, Kiana Jafari, Allie Griffith +4
Across healthcare, agentic artificial intelligence (AI) systems are increasingly promoted as capable of autonomous action, yet in practice they currently operate under near-total h…
cs.CL2025
ASTPrompter: Preference-Aligned Automated Language Model Red-Teaming to Generate Low-Perplexity Unsafe Prompts
Amelia F. Hardy, Houjun Liu, Allie Griffith +3
Existing LLM red-teaming approaches prioritize high attack success rate, often resulting in high-perplexity prompts. This focus overlooks low-perplexity attacks that are more diffi…
cs.AI2024
More than Marketing? On the Information Value of AI Benchmarks for Practitioners
Amelia Hardy, Anka Reuel, Kiana Jafari Meimandi +6
Public AI benchmark results are widely broadcast by model developers as indicators of model quality within a growing and competitive market. However, these advertised scores do not…