2 citations · 2 across the 1 of their papers we have counts for
4 papers
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
Jan Batzner, Sree Harsha Nelaturu, Damian Stachura +45
AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First…
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting
Avijit Ghosh, Anka Reuel, Jenny Chim +45
AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers can…
Multi-Modal Framing Analysis of News
Arnav Arora, Srishti Yadav, Maria Antoniak +2
Automated frame analysis of political communication is a popular task in computational social science that is used to study how authors select aspects of a topic to frame its recep…
Survey of Cultural Awareness in Language Models: Text and Beyond
Siddhesh Pawar, Junyeong Park, Jiho Jin +7
Large-scale deployment of large language models (LLMs) in various applications, such as chatbots and virtual assistants, requires LLMs to be culturally sensitive to the user to ens…