10 papers
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
Xiaomin Li, Yuexing Hao, Jianheng Hou +90
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and inter…
PALMs: Using Multi Construct-Grounded Rationales for Modeling Population Preferences in LLMs
Priyanka Dey, Brihi Joshi, Preyashi Poddar +2
Large language models are being extensively used to simulate individual user behavior, yet faithfully representing a population requires capturing the systematic variation in value…
Translating the Untranslatable: An Operationalizable Ontology for Untranslatability
Jacob Bremerman, Brihi Joshi, Hirona Arai +2
Untranslatability, cases where meaning cannot be directly preserved across languages, is well-studied in linguistics but underexplored in NLP. As machine translation (MT) systems i…
Rigorous Interpretation Is a Form of Evaluation
Isabelle Lee, Emmy Liu, Cathy Jiao +4
Current machine learning models are evaluated through behavioral snapshots, with benchmark accuracies, win rates and outcome-based metrics. Model explanations and evaluations, howe…
Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
Keyu He, Tejas Srinivasan, Brihi Joshi +3
When people query Vision-Language Models (VLMs) but cannot see the accompanying visual context (e.g. for blind and low-vision users), augmenting VLM predictions with natural langua…
PrimeX: A Dataset of Worldview, Opinion, and Explanation
Rik Koncel-Kedziorski, Brihi Joshi, Tim Paek
As the adoption of language models advances, so does the need to better represent individual users to the model. Are there aspects of an individual's belief system that a language…