activity
20242026
collaborators

11 papers

cs.LG2026

CAREBench: A Child-Safety Risk Benchmark for Language Models

Kaavya Krishna-Kumar, Elaine Lau, Vaughn Robinson +6

How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm? Existing child safety evaluations focus on child sexual abuse…

cs.HC2026

"ChatGPT, help me draft a breakup text": The Covert Triad and Articulation Labor in AI-Assisted Romantic Communication

Skyler Wang, Isabella Luppi

Generative artificial intelligence (AI) has begun infiltrating the most ordinary domains of romantic life -- drafting apologies, softening reproaches, and decoding a partner's ambi…

cs.AI2026

SCRuB: Social Concept Reasoning under Rubric-Based Evaluation

Jamelle Watson-Daniels, Himaghna Bhattacharjee, Skyler Wang +11

While many studies of Large Language Model (LLM) reasoning capabilities emphasize mathematical or technical tasks, few address reasoning about social concepts: the abstract ideas s…

cs.LG2026

The Pragmatic Frames of Spurious Correlations in Machine Learning: Interpreting How and Why They Matter

Samuel J. Bell, Skyler Wang

Learning correlations from data forms the foundation of today's machine learning (ML) and artificial intelligence research. While contemporary methods enable the automatic discover…

cs.AI2026

BankerToolBench: Evaluating AI Agents in End-to-End Investment Banking Workflows

Elaine Lau, Markus Dücker, Ronak Chaudhary +24

Existing AI benchmarks lack the fidelity to assess economically meaningful progress on professional workflows. To evaluate frontier AI agents in a high-value, labor-intensive profe…

cs.CL2025

Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages

Omnilingual ASR team, Gil Keren, Artyom Kozhevnikov +30

Automatic speech recognition (ASR) has advanced in high-resource languages, but most of the world's 7,000+ languages remain unsupported, leaving thousands of long-tail languages be…