5 papers
RADAR: Rubric-Aware Dependency and Redundancy Analysis for LLM-as-Judge Evaluation
Divyansh Singh, Reza Davari, Afra Mashhadi
Rubric-based LLM-as-judge pipelines often assume that evaluation criteria provide independent signals. In practice, however, criteria can be behaviorally coupled: improving one cri…
Report on CHIIR 2026 Workshop on Generative AI and Academic Search (GAI&AS)
Yifan Liu, Jaime Arguello, Orland Hoeber +30
This report summarizes the CHIIR 2026 Workshop on Generative AI and Academic Search (GAI\&AS), which examined how GenAI is reshaping academic search systems and research practices.…
Remembering Unequally: Global and Disciplinary Bias in LLM Reconstruction of Scholarly Coauthor Lists
Ghazal Kalhor, Afra Mashhadi
Ongoing breakthroughs in large language models (LLMs) are reshaping scholarly search and discovery interfaces. While these systems offer new possibilities for navigating scientific…
ToxiTwitch: Toward Emote-Aware Hybrid Moderation for Live Streaming Platforms
Baktash Ansari, Elias Martin, Afra Mashhadi
The rapid growth of live-streaming platforms such as Twitch has introduced complex challenges in moderating toxic behavior. Traditional moderation approaches, such as human annotat…
Comparing Fairness of Generative Mobility Models
Daniel Wang, Jack McFarland, Afra Mashhadi +1
This work examines the fairness of generative mobility models, addressing the often overlooked dimension of equity in model performance across geographic regions. Predictive models…