From the 2 of 7 linked papers with an AI index.
6 papers
Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination
Qiyao Liang, Risto Miikkulainen, Ila Fiete
Language models draw on two knowledge sources: facts baked into weights (parametric memory, PM) and information in context (working memory, WM). We study two mechanistically distin…
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
Xin Qiu, Yulu Gan, Conor F. Hayes +6
The paper shows that evolution strategies can successfully fine‑tune billion‑parameter large language models without backpropagation, outperforming reinforcement learning in stabil…
The Blessing of Dimensionality in LLM Fine-tuning: A Variance-Curvature Perspective
Qiyao Liang, Jinyeop Song, Yizhou Liu +4
Weight-perturbation evolution strategies (ES) can fine-tune billion-parameter language models with surprisingly small populations (e.g., ), contradicting classical…
Modular connectivity in neural networks emerges from Poisson noise-motivated regularisation, and promotes robustness and compositional generalisation
Daoyuan Qian, Qiyao Liang, Ila Fiete
Circuits in the brain commonly exhibit modular architectures that factorise complex tasks, resulting in the ability to compositionally generalise and reduce catastrophic forgetting…
Compositional Generalization via Forced Rendering of Disentangled Latents
Qiyao Liang, Daoyuan Qian, Liu Ziyin +1
Composition-the ability to generate myriad variations from finite means-is believed to underlie powerful generalization. However, compositional generalization remains a key challen…
How Diffusion Models Learn to Factorize and Compose
Qiyao Liang, Ziming Liu, Mitchell Ostrow +1
Diffusion models are capable of generating photo-realistic images that combine elements which likely do not appear together in the training set, demonstrating the ability to \texti…