3 citations · 6 across the 4 of their papers we have counts for
8 papers
The Llama 4 Herd: Architecture, Training, Evaluation, and Deployment Notes
Redacted by arXiv
This document consolidates publicly reported technical details about Metas Llama 4 model family. It summarizes (i) released variants (Scout and Maverick) and the broader herd conte…
Do LLMs dream of elephants (when told not to)? Latent concept association and associative memory in transformers
Yibo Jiang, Goutham Rajendran, Pradeep Ravikumar +1
Large Language Models (LLMs) have the capacity to store and recall facts. Through experimentation with open-source models, we observe that this ability to retrieve facts can be eas…
Efficient Certificates of Anti-Concentration Beyond Gaussians
Ainesh Bakshi, Pravesh Kothari, Goutham Rajendran +2
A set of high dimensional points in isotropic position is said to be -anti concentrated if for every direction , the fraction of poin…
On the Origins of Linear Representations in Large Language Models
Yibo Jiang, Goutham Rajendran, Pradeep Ravikumar +2
Recent works have argued that high-level semantic concepts are encoded "linearly" in the representation space of large language models. In this work, we study the origins of such l…
Learning Interpretable Concepts: Unifying Causal Representation Learning and Foundation Models
Goutham Rajendran, Simon Buchholz, Bryon Aragam +2
To build intelligent machine learning systems, there are two broad approaches. One approach is to build inherently interpretable models, as endeavored by the growing field of causa…
An Interventional Perspective on Identifiability in Gaussian LTI Systems with Independent Component Analysis
Goutham Rajendran, Patrik Reizinger, Wieland Brendel +1
We investigate the relationship between system identification and intervention design in dynamical systems. While previous research demonstrated how identifiable representation lea…