6 papers
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)
Ro Encarnación, Tina Behzad, Emma Lurie +1
Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet most evaluations rely on a sing…
Do LLMs Ask the Right Questions? Evaluating GPT-Generated Surveys as Instruments for Measuring Social Attitudes
Tina Behzad, Wenbo Li, Reuben Kline +1
Understanding human beliefs and social attitudes often relies on carefully designed survey instruments. Recent work has suggested that large language models (LLMs) could automate p…
An External Fairness Evaluation of LinkedIn Talent Search
Tina Behzad, Siddartha Devic, Vatsal Sharan +2
We conduct an independent, third-party audit for bias of LinkedIn's Talent Search ranking system, focusing on potential ranking bias across two attributes: gender and race. To do s…
Beyond Predictions: A Study of AI Strength and Weakness Transparency Communication on Human-AI Collaboration
Tina Behzad, Nikolos Gurney, Ning Wang +1
The promise of human-AI teaming lies in humans and AI working together to achieve performance levels neither could accomplish alone. Effective communication between AI and humans i…
FairPlay: A Collaborative Approach to Mitigate Bias in Datasets for Improved AI Fairness
Tina Behzad, Mithilesh Kumar Singh, Anthony J. Ripa +1
The issue of fairness in decision-making is a critical one, especially given the variety of stakeholder demands for differing and mutually incompatible versions of fairness. Adopti…
Reconciling Predictive Multiplicity in Practice
Tina Behzad, SÃlvia Casacuberta, Emily Ruth Diana +1
Many machine learning applications predict individual probabilities, such as the likelihood that a person develops a particular illness. Since these probabilities are unknown, a ke…