activity
20242026
collaborators

10 papers

cs.AI2026

Can AI Agents Synthesize Scientific Conclusions?

Hayoung Jung, Pedro Viana Diniz, José Reinaldo Corrêa Roveda +5

Scientific AI agents increasingly retrieve evidence, reason across sources, and synthesize conclusions used in consequential decisions. Yet, their ability to do so in high-stakes d…

cs.CV2026

WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark

Yida Yin, Harish Krishnakumar, Chung Peng Lee +9

In real-world applications, models are expected to perform reliably across diverse settings. Yet, many existing multimodal benchmarks expand task types without capturing the visual…

cs.CY2026

Measuring Validity in LLM-based Resume Screening

Jane Castleman, Zeyu Shen, Blossom Metevier +2

Resume screening is perceived as a particularly suitable task for LLMs given their ability to analyze natural language; thus many entities rely on general purpose LLMs without furt…

cs.LG2026

The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety

Max Springer, Chung Peng Lee, Blossom Metevier +5

Fine-tuning aligned language models on benign tasks unpredictably degrades safety guardrails, even when training data contains no harmful content and developers have no adversarial…

cs.LG2025

FrontierCS: Evolving Challenges for Evolving Intelligence

Qiuyang Mang, Wenhao Chai, Zhifei Li +48

We introduce FrontierCS, a benchmark of 156 open-ended problems across diverse areas of computer science, designed and reviewed by experts, including CS PhDs and top-tier competiti…

cs.CY2025

An External Fairness Evaluation of LinkedIn Talent Search

Tina Behzad, Siddartha Devic, Vatsal Sharan +2

We conduct an independent, third-party audit for bias of LinkedIn's Talent Search ranking system, focusing on potential ranking bias across two attributes: gender and race. To do s…