4 papers
Characterizing Cultural Localization in AI-Generated Stories
Shaily Bhatt, Supriti Vijay, Jeremiah Milbauer +1
The global use of artificial intelligence has increased interest in assessing the ability to generate culturally localized content, including stories. Cultural localization in stor…
Beyond Text: Characterizing Domain Expert Needs in Document Research
Sireesh Gururaja, Nupoor Gandhi, Jeremiah Milbauer +1
Working with documents is a key part of almost any knowledge work, from contextualizing research in a literature review to reviewing legal precedent. Recently, as their capabilitie…
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
Stereotype or Personalization? User Identity Biases Chatbot Recommendations
Anjali Kantharuban, Jeremiah Milbauer, Maarten Sap +2
While personalized recommendations are often desired by users, it can be difficult in practice to distinguish cases of bias from cases of personalization: we find that models gener…