activity
20242026
collaborators

8 papers

cs.CL2026

Diffusion-Inspired Masked Fine-Tuning for Knowledge Injection in Autoregressive LLMs

Xu Pan, Ely Hahami, Jingxuan Fan +2

Large language models (LLMs) are often used in environments where facts evolve, yet factual knowledge updates via fine-tuning on unstructured text often suffer from 1) reliance on…

cond-mat.dis-nn2025

Simplified derivations for high-dimensional convex learning problems

David G. Clark, Haim Sompolinsky

Statistical-physics calculations in machine learning and theoretical neuroscience often involve lengthy derivations that obscure physical interpretation. Here, we give concise, non…

q-bio.NC2025

Unraveling the geometry of visual relational reasoning

Jiaqi Shang, Gabriel Kreiman, Haim Sompolinsky

Humans readily generalize abstract relations, such as recognizing "constant" in shape or color, whereas neural networks struggle, limiting their flexible reasoning. To investigate…

cs.LG2025

Connecting NTK and NNGP: A Unified Theoretical Framework for Wide Neural Network Learning Dynamics

Yehonatan Avidan, Qianyi Li, Haim Sompolinsky

Artificial neural networks have revolutionized machine learning in recent years, but a complete theoretical framework for their learning process is still lacking. Substantial advan…

cs.CL2025

Memorization and Knowledge Injection in Gated LLMs

Xu Pan, Ely Hahami, Zechen Zhang +1

Large Language Models (LLMs) currently struggle to sequentially add new memories and integrate new knowledge. These limitations contrast with the human ability to continuously lear…

cs.LG2025

When narrower is better: the narrow width limit of Bayesian parallel branching neural networks

Zechen Zhang, Haim Sompolinsky

The infinite width limit of random neural networks is known to result in Neural Networks as Gaussian Process (NNGP) (Lee et al. (2018)), characterized by task-independent kernels.…