Publications (9)
A Roadmap to Pluralistic Alignment
Taylor Sorensen, Jared Moore, Jillian Fisher +9
With increased power and prevalence of AI systems, it is ever more critical that AI systems are designed to serve all, i.e., people with diverse values and perspectives. However, a…
A Roadmap to Impactful Pluralistic Alignment Research
Elinor Poole-Dayan, Jillian Fisher, Atoosa Kasirzadeh +3
Pluralistic value alignment---the goal of building AI systems that represent and serve diverse human values and perspectives---has emerged as an active research agenda. Yet, there'…
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
Xiaomin Li, Yuexing Hao, Jianheng Hou +90
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and inter…
DracoGPT: Extracting Visualization Design Preferences from Large Language Models
Huichen Will Wang, Mitchell Gordon, Leilani Battle +1
Trained on vast corpora, Large Language Models (LLMs) have the potential to encode visualization design knowledge and best practices. However, if they fail to do so, they might pro…
StyleRemix: Interpretable Authorship Obfuscation via Distillation and Perturbation of Style Elements
Jillian Fisher, Skyler Hallinan, Ximing Lu +3
Authorship obfuscation, rewriting a text to intentionally obscure the identity of the author, is an important but challenging task. Current methods using large language models (LLM…
Localizing Paragraph Memorization in Language Models
Niklas Stoehr, Mitchell Gordon, Chiyuan Zhang +1
Can we localize the weights and mechanisms used by a language model to memorize and recite entire paragraphs of its training data? In this paper, we show that while memorization is…