4 papers
Associative memory inspires improvements for in-context learning using a novel attention residual stream architecture
Thomas F Burns, Tomoki Fukai, Christopher J Earls
Large language models (LLMs) demonstrate an impressive ability to utilise information within the context of their input sequences to appropriately respond to data unseen by the LLM…
Density estimation with LLMs: a geometric investigation of in-context learning trajectories
Toni J. B. Liu, Nicolas Boullé, Raphaël Sarfati +1
Large language models (LLMs) demonstrate remarkable emergent abilities to perform in-context learning across various tasks, including time series forecasting. This work investigate…
Lines of Thought in Large Language Models
Raphaël Sarfati, Toni J. B. Liu, Nicolas Boullé +1
Large Language Models achieve next-token prediction by transporting a vectorized piece of text (prompt) across an accompanying embedding space under the action of successive transf…
LLMs learn governing principles of dynamical systems, revealing an in-context neural scaling law
Toni J. B. Liu, Nicolas Boullé, Raphaël Sarfati +1
Pretrained large language models (LLMs) are surprisingly effective at performing zero-shot tasks, including time-series forecasting. However, understanding the mechanisms behind su…