2 papers
cs.LG2026
What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus
Edward Lue Chee Lip, Boden Moraski, Tim Knappe +4
Frontier benchmarks need tasks that current models cannot solve. But a task that no model solves is not automatically a hard task. The same zero pass rate can come from a real capa…
cs.HC2026
Anthropomorphism in AI Companion Communities: Age, Gender, and Emotional Correlates
Afia Mubashir, Boden Moraski, Stephanie Choi +1
Artificial intelligence (AI) systems are increasingly integrated into daily life, with millions now using AI chatbots built on Large Language Models (LLMs) for companionship. Both…