10 papers
Can AI Agents Synthesize Scientific Conclusions?
Hayoung Jung, Pedro Viana Diniz, José Reinaldo Corrêa Roveda +5
Scientific AI agents increasingly retrieve evidence, reason across sources, and synthesize conclusions used in consequential decisions. Yet, their ability to do so in high-stakes d…
WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark
Yida Yin, Harish Krishnakumar, Chung Peng Lee +9
In real-world applications, models are expected to perform reliably across diverse settings. Yet, many existing multimodal benchmarks expand task types without capturing the visual…
Measuring Validity in LLM-based Resume Screening
Jane Castleman, Zeyu Shen, Blossom Metevier +2
Resume screening is perceived as a particularly suitable task for LLMs given their ability to analyze natural language; thus many entities rely on general purpose LLMs without furt…
The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety
Max Springer, Chung Peng Lee, Blossom Metevier +5
Fine-tuning aligned language models on benign tasks unpredictably degrades safety guardrails, even when training data contains no harmful content and developers have no adversarial…
FrontierCS: Evolving Challenges for Evolving Intelligence
Qiuyang Mang, Wenhao Chai, Zhifei Li +48
We introduce FrontierCS, a benchmark of 156 open-ended problems across diverse areas of computer science, designed and reviewed by experts, including CS PhDs and top-tier competiti…
An External Fairness Evaluation of LinkedIn Talent Search
Tina Behzad, Siddartha Devic, Vatsal Sharan +2
We conduct an independent, third-party audit for bias of LinkedIn's Talent Search ranking system, focusing on potential ranking bias across two attributes: gender and race. To do s…