8 papers
Diffusion-Inspired Masked Fine-Tuning for Knowledge Injection in Autoregressive LLMs
Xu Pan, Ely Hahami, Jingxuan Fan +2
Large language models (LLMs) are often used in environments where facts evolve, yet factual knowledge updates via fine-tuning on unstructured text often suffer from 1) reliance on…
Simplified derivations for high-dimensional convex learning problems
David G. Clark, Haim Sompolinsky
Statistical-physics calculations in machine learning and theoretical neuroscience often involve lengthy derivations that obscure physical interpretation. Here, we give concise, non…
Unraveling the geometry of visual relational reasoning
Jiaqi Shang, Gabriel Kreiman, Haim Sompolinsky
Humans readily generalize abstract relations, such as recognizing "constant" in shape or color, whereas neural networks struggle, limiting their flexible reasoning. To investigate…
Connecting NTK and NNGP: A Unified Theoretical Framework for Wide Neural Network Learning Dynamics
Yehonatan Avidan, Qianyi Li, Haim Sompolinsky
Artificial neural networks have revolutionized machine learning in recent years, but a complete theoretical framework for their learning process is still lacking. Substantial advan…
Memorization and Knowledge Injection in Gated LLMs
Xu Pan, Ely Hahami, Zechen Zhang +1
Large Language Models (LLMs) currently struggle to sequentially add new memories and integrate new knowledge. These limitations contrast with the human ability to continuously lear…
When narrower is better: the narrow width limit of Bayesian parallel branching neural networks
Zechen Zhang, Haim Sompolinsky
The infinite width limit of random neural networks is known to result in Neural Networks as Gaussian Process (NNGP) (Lee et al. (2018)), characterized by task-independent kernels.…