6 papers
Representational Similarity and Model Behavior in Multi-Agent Interaction
Yujin Potter, Seun Eisape, Shiyang Lai +6
Researchers have shown that neural similarity among humans predicts social closeness and cooperative success, whereas innovation often emerges from interactions among dissimilar in…
Escaping Plato's Cave: JAM for Aligning Independently Trained Vision and Language Models
Lauren Hyoseo Yoon, Yisong Yue, Been Kim
Independently trained vision and language models inhabit disjoint representational spaces, shaped by their respective modalities, objectives, and architectures. The Platonic Repres…
Neologism Learning for Controllability and Self-Verbalization
John Hewitt, Oyvind Tafjord, Robert Geirhos +1
Humans invent new words when there is a rising demand for a new useful concept (e.g., doomscrolling). We explore and validate a similar idea in our communication with LLMs: introdu…
How many classes do we need to see for novel class discovery?
Akanksha Sarkar, Been Kim, Jennifer J. Sun
Novel class discovery is essential for ML models to adapt to evolving real-world data, with applications ranging from scientific discovery to robotics. However, these datasets cont…
Because we have LLMs, we Can and Should Pursue Agentic Interpretability
Been Kim, John Hewitt, Neel Nanda +2
The era of Large Language Models (LLMs) presents a new opportunity for interpretability--agentic interpretability: a multi-turn conversation with an LLM wherein the LLM proactively…
We Can't Understand AI Using our Existing Vocabulary
John Hewitt, Robert Geirhos, Been Kim
This position paper argues that, in order to understand AI, we cannot rely on our existing vocabulary of human words. Instead, we should strive to develop neologisms: new words tha…