2 papers
cs.CY2025
Defending Compute Thresholds Against Legal Loopholes
Matteo Pistillo, Pablo Villalobos
Existing legal frameworks on AI rely on training compute thresholds as a proxy to identify potentially-dangerous AI models and trigger increased regulatory attention. In the United…
cs.LG2024
Will we run out of data? Limits of LLM scaling based on human-generated data
Pablo Villalobos, Anson Ho, Jaime Sevilla +3
We investigate the potential constraints on LLM scaling posed by the availability of public human-generated text data. We forecast the growing demand for training data based on cur…