collaborators

6 papers

cs.CL2026

Representational Similarity and Model Behavior in Multi-Agent Interaction

Yujin Potter, Seun Eisape, Shiyang Lai +6

Researchers have shown that neural similarity among humans predicts social closeness and cooperative success, whereas innovation often emerges from interactions among dissimilar in…

cs.LG2026

Escaping Plato's Cave: JAM for Aligning Independently Trained Vision and Language Models

Lauren Hyoseo Yoon, Yisong Yue, Been Kim

Independently trained vision and language models inhabit disjoint representational spaces, shaped by their respective modalities, objectives, and architectures. The Platonic Repres…

cs.CL2025

Neologism Learning for Controllability and Self-Verbalization

John Hewitt, Oyvind Tafjord, Robert Geirhos +1

Humans invent new words when there is a rising demand for a new useful concept (e.g., doomscrolling). We explore and validate a similar idea in our communication with LLMs: introdu…

cs.LG2025

How many classes do we need to see for novel class discovery?

Akanksha Sarkar, Been Kim, Jennifer J. Sun

Novel class discovery is essential for ML models to adapt to evolving real-world data, with applications ranging from scientific discovery to robotics. However, these datasets cont…

cs.AI2025

Because we have LLMs, we Can and Should Pursue Agentic Interpretability

Been Kim, John Hewitt, Neel Nanda +2

The era of Large Language Models (LLMs) presents a new opportunity for interpretability--agentic interpretability: a multi-turn conversation with an LLM wherein the LLM proactively…

cs.CL2025

We Can't Understand AI Using our Existing Vocabulary

John Hewitt, Robert Geirhos, Been Kim

This position paper argues that, in order to understand AI, we cannot rely on our existing vocabulary of human words. Instead, we should strive to develop neologisms: new words tha…