11 papers
CAREBench: A Child-Safety Risk Benchmark for Language Models
Kaavya Krishna-Kumar, Elaine Lau, Vaughn Robinson +6
How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm? Existing child safety evaluations focus on child sexual abuse…
"ChatGPT, help me draft a breakup text": The Covert Triad and Articulation Labor in AI-Assisted Romantic Communication
Skyler Wang, Isabella Luppi
Generative artificial intelligence (AI) has begun infiltrating the most ordinary domains of romantic life -- drafting apologies, softening reproaches, and decoding a partner's ambi…
SCRuB: Social Concept Reasoning under Rubric-Based Evaluation
Jamelle Watson-Daniels, Himaghna Bhattacharjee, Skyler Wang +11
While many studies of Large Language Model (LLM) reasoning capabilities emphasize mathematical or technical tasks, few address reasoning about social concepts: the abstract ideas s…
The Pragmatic Frames of Spurious Correlations in Machine Learning: Interpreting How and Why They Matter
Samuel J. Bell, Skyler Wang
Learning correlations from data forms the foundation of today's machine learning (ML) and artificial intelligence research. While contemporary methods enable the automatic discover…
BankerToolBench: Evaluating AI Agents in End-to-End Investment Banking Workflows
Elaine Lau, Markus Dücker, Ronak Chaudhary +24
Existing AI benchmarks lack the fidelity to assess economically meaningful progress on professional workflows. To evaluate frontier AI agents in a high-value, labor-intensive profe…
Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages
Omnilingual ASR team, Gil Keren, Artyom Kozhevnikov +30
Automatic speech recognition (ASR) has advanced in high-resource languages, but most of the world's 7,000+ languages remain unsupported, leaving thousands of long-tail languages be…