activity
20242026
collaborators

7 papers

cs.MA2026

Emergence of Biased Consensus in Multi-Agent LLM Debates

Maya Okawa

Multi-agent LLM debates achieve strong performance on decision-making tasks as well as problem-solving benchmarks, yet their safety and fairness risks remain poorly understood. Not…

cs.CL2026

Emergence of Hierarchical Emotion Organization in Large Language Models

Maya Okawa, Bo Zhao, Eric J. Bigelow +4

As large language models (LLMs) increasingly power conversational agents, understanding how they model users' emotional states is critical for ethical deployment. Inspired by emoti…

cs.LG2025

Swing-by Dynamics in Concept Learning and Compositional Generalization

Yongyi Yang, Core Francisco Park, Ekdeep Singh Lubana +3

Prior work has shown that text-conditioned diffusion models can learn to identify and manipulate primitive concepts underlying a compositional data-generating process, enabling gen…

cs.LG2025

Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task

Maya Okawa, Ekdeep Singh Lubana, Robert P. Dick +1

Modern generative models exhibit unprecedented capabilities to generate extremely realistic data. However, given the inherent compositionality of the real world, reliable use of th…

cs.LG2025

Representation Shattering in Transformers: A Synthetic Study with Knowledge Editing

Kento Nishi, Rahul Ramesh, Maya Okawa +3

Knowledge Editing (KE) algorithms alter models' weights to perform targeted updates to incorrect, outdated, or otherwise unwanted factual associations. However, recent work has sho…

cs.CL2025

ICLR: In-Context Learning of Representations

Core Francisco Park, Andrew Lee, Ekdeep Singh Lubana +5

Recent work has demonstrated that semantics specified by pretraining data influence how representations of different concepts are organized in a large language model (LLM). However…