2 papers
cs.CL2026
Invisible failures in human-AI interactions
Christopher Potts, Moritz Sudhof
AI systems fail silently far more often than they fail visibly. In an analysis of 100K human-AI interactions from the WildChat dataset, we find that 79% of AI failures are invisibl…
cs.CL2026
Structured Prompts Improve Evaluation of Language Models
Asad Aali, Muhammad Ahmed Mohsin, Vasiliki Bikia +15
As language models (LMs) are increasingly adopted across domains, high-quality benchmarking frameworks are essential for guiding deployment decisions. In practice, however, framewo…