17 citations · 33 across the 12 of their papers we have counts for
4 papers · 1 filter
AdaSplash-2: Faster Differentiable Sparse Attention
Nuno Gonçalves, Hugo Pitorro, Vlad Niculae +4
Sparse attention has been proposed as a way to alleviate the quadratic cost of transformers, a central bottleneck in long-context training. A promising line of work is -entmax a…
Adapting Time Series Foundation Models through Data Mixtures
Thomas L. Lee, Edoardo M. Ponti, Amos Storkey
Time series foundation models (TSFMs) have become increasingly popular for zero-shot forecasting. However, for a new time series domain not fully covered by the pretraining set, pe…
Is Information Density Uniform when Utterances are Grounded on Perception and Discourse?
Matteo Gay, Coleman Haley, Mario Giulianelli +1
The Uniform Information Density (UID) hypothesis posits that speakers are subject to a communicative pressure to distribute information evenly within utterances, minimising surpris…
Self-Improving World Modelling with Latent Actions
Yifu Qiu, Zheng Zhao, Waylon Li +4
Internal modelling of the world -- predicting transitions between previous states and next states under actions -- is essential to reasoning and planning for LLMs and V…