5 citations · 6 across the 4 of their papers we have counts for
4 papers
LeRAAT: LLM-Enabled Real-Time Aviation Advisory Tool
Marc R. Schlichting, Vale Rasmussen, Heba Alazzeh +5
In aviation emergencies, high-stakes decisions must be made in an instant. Pilots rely on quick access to precise, context-specific information -- an area where emerging tools like…
More than Marketing? On the Information Value of AI Benchmarks for Practitioners
Amelia Hardy, Anka Reuel, Kiana Jafari Meimandi +6
Public AI benchmark results are widely broadcast by model developers as indicators of model quality within a growing and competitive market. However, these advertised scores do not…
BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices
Anka Reuel, Amelia Hardy, Chandler Smith +3
AI models are increasingly prevalent in high-stakes environments, necessitating thorough assessment of their capabilities and risks. Benchmarks are popular for measuring these attr…
ASTPrompter: Preference-Aligned Automated Language Model Red-Teaming to Generate Low-Perplexity Unsafe Prompts
Amelia F. Hardy, Houjun Liu, Allie Griffith +3
Existing LLM red-teaming approaches prioritize high attack success rate, often resulting in high-perplexity prompts. This focus overlooks low-perplexity attacks that are more diffi…