Human-like object concept representations emerge naturally in multimodal large language models
arXiv:2407.01067 · doi:10.1038/s42256-025-01049-z
Abstract
Understanding how humans conceptualize and categorize natural objects offers critical insights into perception and cognition. With the advent of Large Language Models (LLMs), a key question arises: can these models develop human-like object representations from linguistic and multimodal data? In this study, we combined behavioral and neuroimaging analyses to explore the relationship between object concept representations in LLMs and human cognition. We collected 4.7 million triplet judgments from LLMs and Multimodal LLMs (MLLMs) to derive low-dimensional embeddings that capture the similarity structure of 1,854 natural objects. The resulting 66-dimensional embeddings were stable, predictive, and exhibited semantic clustering similar to human mental representations. Remarkably, the dimensions underlying these embeddings were interpretable, suggesting that LLMs and MLLMs develop human-like conceptual representations of objects. Further analysis showed strong alignment between model embeddings and neural activity patterns in brain regions such as EBA, PPA, RSC, and FFA. This provides compelling evidence that the object representations in LLMs, while not identical to human ones, share fundamental similarities that reflect key aspects of human conceptual knowledge. Our findings advance the understanding of machine intelligence and inform the development of more human-like artificial cognitive systems.
Published on Nature Machine Intelligence
References in corpus (17)
- Learning Transferable Visual Models From Natural Language Supervision
- A Survey on Multimodal Large Language Models
- Using cognitive psychology to understand GPT-3
- Understanding the Role of Individual Units in a Deep Neural Network
- Thinking Fast and Slow in Large Language Models
- Human-Like Intuitive Behavior and Reasoning Biases Emerged in Language Models -- and Disappeared in GPT-4
- From Attribution Maps to Human-Understandable Explanations through Concept Relevance Propagation
- Acquisition of Chess Knowledge in AlphaZero
- The low-rank hypothesis of complex systems
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Capturing human categorization of natural images at scale by combining deep networks and cognitive models
- Large Language Models Understand and Can be Enhanced by Emotional Stimuli
- Visual representations in the human brain are aligned with large language models
- Learning as the Unsupervised Alignment of Conceptual Systems
- Gromov-Wasserstein unsupervised alignment reveals structural correspondences between the color similarity structures of humans and large language models
- CoCoG: Controllable Visual Stimuli Generation based on Human Concept Representations
- Comparing Rationality Between Large Language Models and Humans: Insights and Open Questions