papers

Publications (52)

cs.SD2022

Multi-instrument Music Synthesis with Spectrogram Diffusion

Curtis Hawthorne, Ian Simon, Adam Roberts +4

An ideal music synthesizer should be both interactive and expressive, generating high-fidelity audio in realtime for arbitrary combinations of instruments and notes. Recent neural…

cs.SD2019

The Bach Doodle: Approachable music composition with machine learning at scale

Cheng-Zhi Anna Huang, Curtis Hawthorne, Adam Roberts +4

To make music composition more approachable, we designed the first AI-powered Google Doodle, the Bach Doodle, where users can create their own melody and have it harmonized by a ma…

cs.AI2023

The Flan Collection: Designing Data and Methods for Effective Instruction Tuning

Shayne Longpre, Le Hou, Tu Vu +8

We study the design decisions of publicly available instruction tuning methods, and break down the development of Flan 2022 (Chung et al., 2022). Through careful ablation studies o…

cs.CR2021

Extracting Training Data from Large Language Models

Nicholas Carlini, Florian Tramer, Eric Wallace +9

It has become common to publish large (billion parameter) language models that have been trained on private datasets. This paper demonstrates that in such settings, an adversary ca…

cs.LG2023

Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Colin Raffel, Noam Shazeer, Adam Roberts +6

Transfer learning, where a model is first pre-trained on a data-rich task before being fine-tuned on a downstream task, has emerged as a powerful technique in natural language proc…

physics.ins-det2026

The Darkside-20k Data Acquisition System

Fabio Acerbi, Pushparaj Adhikari, Paolo Agnes +289

The paper describes the data acquisition (DAQ) system for the DarkSide‑20k liquid‑argon dark‑matter detector, which uses continuous, trigger‑less digitisation of 2720 SiPM channels…

#dark matter detection#liquid argon detector#data acquisition#triggerless readout
cs.CL2020

How Much Knowledge Can You Pack Into the Parameters of a Language Model?

Adam Roberts, Colin Raffel, Noam Shazeer

It has recently been observed that neural language models trained on unstructured text can implicitly store and retrieve knowledge using natural language queries. In this short pap…

cs.HC2025

ReaLJam: Real-Time Human-AI Music Jamming with Reinforcement Learning-Tuned Transformers

Alexander Scarlatos, Yusong Wu, Ian Simon +5

Recent advances in generative artificial intelligence (AI) have created models capable of high-quality musical content generation. However, little consideration is given to how to…

cs.CL2023

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

BigScience Workshop, :, Teven Le Scao +391

Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to wi…

cs.LG2021

Do Transformer Modifications Transfer Across Implementations and Applications?

Sharan Narang, Hyung Won Chung, Yi Tay +13

The research community has proposed copious modifications to the Transformer architecture since it was introduced over three years ago, relatively few of which have seen widespread…

cond-mat.mtrl-sci2011

Response of graphene to femtosecond high-intensity laser irradiation

Adam Roberts, Daniel Cormode, Collin Reynolds +3

We study the response of graphene to high-intensity 10^11-10^12 Wcm^-2, 50-femtosecond laser pulse excitation. We establish that graphene has a fairly high (~3\times10^12Wcm^-2) si…

physics.ins-det2023

ARIADNE+: Large scale demonstration of fast optical readout for dual-phase LArTPCs at the CERN Neutrino Platform

Adam Lowe, Pablo Amedo, Diego González-Díaz +12

Optical readout of large scale dual-phase liquid Argon TPCs is an attractive alternative to charge readout and has been successfully demonstrated on a 2x2m active region within the…

cs.CL2023

Crosslingual Generalization through Multitask Finetuning

Niklas Muennighoff, Thomas Wang, Lintang Sutawika +16

Multitask prompted finetuning (MTF) has been shown to help large language models generalize to new tasks in a zero-shot setting, but so far explorations of MTF have focused on Engl…

cs.CL2022

ByT5: Towards a token-free future with pre-trained byte-to-byte models

Linting Xue, Aditya Barua, Noah Constant +5

Most widely-used pre-trained language models operate on sequences of tokens corresponding to word or subword units. By comparison, token-free models that operate directly on raw te…

cs.CL2024

Gemma: Open Models Based on Gemini Research and Technology

Gemma Team, Thomas Mesnard, Cassidy Hardin +105

This work introduces Gemma, a family of lightweight, state-of-the art open models built from the research and technology used to create Gemini models. Gemma models demonstrate stro…

cs.SD2023

MusicLM: Generating Music From Text

Andrea Agostinelli, Timo I. Denk, Zalán Borsos +10

We introduce MusicLM, a model generating high-fidelity music from text descriptions such as "a calming violin melody backed by a distorted guitar riff". MusicLM casts the process o…

cs.SD2019

Enabling Factorized Piano Music Modeling and Generation with the MAESTRO Dataset

Curtis Hawthorne, Andriy Stasyuk, Adam Roberts +6

Generating musical audio directly with neural networks is notoriously difficult because it requires coherently modeling structure at many different timescales. Fortunately, most mu…

stat.ML2018

Learning a Latent Space of Multitrack Measures

Ian Simon, Adam Roberts, Colin Raffel +3

Discovering and exploring the underlying structure of multi-instrumental music using learning-based approaches remains an open problem. We extend the recent MusicVAE model to repre…

cs.LG2019

Counterpoint by Convolution

Cheng-Zhi Anna Huang, Tim Cooijmans, Adam Roberts +2

Machine learning models of music typically break up the task of composition into a chronological process, composing a piece of music in a single pass from beginning to end. On the…

cs.CL2025

T5Gemma 2: Seeing, Reading, and Understanding Longer

Biao Zhang, Paul Suganthan, Gaël Liu +17

We introduce T5Gemma 2, the next generation of the T5Gemma family of lightweight open encoder-decoder models, featuring strong multilingual, multimodal and long-context capabilitie…

cs.SD2019

Learning to Groove with Inverse Sequence Transformations

Jon Gillick, Adam Roberts, Jesse Engel +2

We explore models for translating abstract musical ideas (scores, rhythms) into expressive performances using Seq2Seq and recurrent Variational Information Bottleneck (VIB) models.…

math.GR2009

On a Quasi-Phan Theorem for Orthogonal Groups

Corneliu Hoffman, Adam Roberts

This paper constructs a presentation for the orthogonal groups. The amalgam is contained in that constructed by Ralf Gramlich Hendrik Van Maldeghem and arXiv:0708.1583 but it's sho…

hep-ex2025

Letter of Intent: The Forward Physics Facility

Luis A. Anchordoqui, John K. Anders, Akitaka Ariga +111

The Forward Physics Facility (FPF) is a proposed extension of the HL-LHC program designed to exploit the unique scientific opportunities offered by the intense flux of high energy…

cs.CY2023

Report of the 1st Workshop on Generative AI and Law

A. Feder Cooper, Katherine Lee, James Grimmelmann +32

This report presents the takeaways of the inaugural Workshop on Generative AI and Law (GenLaw), held in July 2023. A cross-disciplinary group of practitioners and scholars from com…

cs.SD2025

Live Music Models

Lyria Team, Antoine Caillon, Brian McWilliams +33

We introduce a new class of generative models for music called live music models that produce a continuous stream of music in real-time with synchronized user control. We release M…

cs.CL2022

What Language Model Architecture and Pretraining Objective Work Best for Zero-Shot Generalization?

Thomas Wang, Adam Roberts, Daniel Hesslow +5

Large pretrained Transformer language models have been shown to exhibit zero-shot generalization, i.e. they can perform a wide variety of tasks that they were not explicitly traine…

cs.CL2025

EmbeddingGemma: Powerful and Lightweight Text Representations

Henrique Schechter Vera, Sahil Dua, Biao Zhang +86

We introduce EmbeddingGemma, a new lightweight, open text embedding model based on the Gemma 3 language model family. Our innovative training recipe strategically captures knowledg…

cs.LG2022

VeLO: Training Versatile Learned Optimizers by Scaling Up

Luke Metz, James Harrison, C. Daniel Freeman +8

While deep learning models have replaced hand-designed features across many domains, these models are still trained with hand-designed optimizers. In this work, we leverage the sam…

cs.CL2023

Large Language Models Struggle to Learn Long-Tail Knowledge

Nikhil Kandpal, Haikang Deng, Adam Roberts +2

The Internet contains a wealth of knowledge -- from the birthdays of historical figures to tutorials on how to code -- all of which may be learned by language models. However, whil…

cs.SD2023

SingSong: Generating musical accompaniments from singing

Chris Donahue, Antoine Caillon, Adam Roberts +8

We present SingSong, a system that generates instrumental music to accompany input vocals, potentially offering musicians and non-musicians alike an intuitive new way to create mus…

cs.CL2022

PaLM: Scaling Language Modeling with Pathways

Aakanksha Chowdhery, Sharan Narang, Jacob Devlin +64

Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of…

physics.ins-det2020

Optical Readout of the ARIADNE LArTPC using a Timepix3-based Camera

Adam Lowe, Krishanu Majumdar, Konstantinos Mavrokoridis +4

The ARIADNE Experiment, utilising a 1-ton dual-phase Liquid Argon Time Projection Chamber (LArTPC), aims to develop and mature optical readout technology for large scale LAr detect…

cs.LG2019

A Hierarchical Latent Vector Model for Learning Long-Term Structure in Music

Adam Roberts, Jesse Engel, Colin Raffel +2

The Variational Autoencoder (VAE) has proven to be an effective model for producing semantically meaningful latent representations for natural data. However, it has thus far seen l…

cs.CL2021

NeurIPS 2020 EfficientQA Competition: Systems, Analyses and Lessons Learned

Sewon Min, Jordan Boyd-Graber, Chris Alberti +50

We review the EfficientQA competition from NeurIPS 2020. The competition focused on open-domain question answering (QA), where systems take natural language questions as input and…

cs.CL2021

mT5: A massively multilingual pre-trained text-to-text transformer

Linting Xue, Noah Constant, Adam Roberts +5

The recent "Text-to-Text Transfer Transformer" (T5) leveraged a unified text-to-text format and scale to attain state-of-the-art results on a wide variety of English-language NLP t…

cs.LG2017

Latent Constraints: Learning to Generate Conditionally from Unconditional Generative Models

Jesse Engel, Matthew Hoffman, Adam Roberts

Deep generative neural networks have proven effective at both conditional and unconditional modeling of complex data distributions. Conditional generation enables interactive contr…

cs.CL2023

A Pretrainer's Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity

Shayne Longpre, Gregory Yauney, Emily Reif +8

Pretraining is the preliminary and fundamental step in developing capable language models (LM). Despite this, pretraining data design is critically under-documented and often guide…

cs.LG2020

DDSP: Differentiable Digital Signal Processing

Jesse Engel, Lamtharn Hantrakul, Chenjie Gu +1

Most generative models of audio directly generate samples in one of two domains: time or frequency. While sufficient to express any signal, these representations are inefficient, a…

q-bio.GN2012

Quantifying uniformity of mapped reads

Valerie Hower, Richard Starfield, Adam Roberts +1

Summary: We describe a tool for quantifying the uniformity of mapped reads in high-throughput sequencing experiments. Our statistic directly measures the uniformity of both read po…

cs.LG2017

Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders

Jesse Engel, Cinjon Resnick, Adam Roberts +4

Generative models in vision have seen rapid progress due to algorithmic improvements and the availability of high-quality image datasets. In this paper, we offer contributions in b…

cs.CL2022

LaMDA: Language Models for Dialog Applications

Romal Thoppilan, Daniel De Freitas, Jamie Hall +57

We present LaMDA: Language Models for Dialog Applications. LaMDA is a family of Transformer-based neural language models specialized for dialog, which have up to 137B parameters an…

cs.LG2022

Scaling Instruction-Finetuned Language Models

Hyung Won Chung, Le Hou, Shayne Longpre +32

Finetuning language models on a collection of datasets phrased as instructions has been shown to improve model performance and generalization to unseen tasks. In this paper we expl…

cs.CL2020

WT5?! Training Text-to-Text Models to Explain their Predictions

Sharan Narang, Colin Raffel, Katherine Lee +3

Neural networks have recently achieved human-level performance on various challenging natural language processing (NLP) tasks, but it is notoriously difficult to understand why a n…

cs.CL2023

UniMax: Fairer and more Effective Language Sampling for Large-Scale Multilingual Pretraining

Hyung Won Chung, Noah Constant, Xavier Garcia +4

Pretrained multilingual large language models have typically used heuristic temperature-based sampling to balance between different languages. However previous work has not systema…

physics.ins-det2024

Feasibility study of a novel thermal neutron detection system using event mode camera and LYSO scintillation crystal

Tianqi Gao, Mohammad Alsulimane, Sergey Burdin +8

The feasibility study of a new technique for thermal neutron detection using a Timepix3 camera (TPX3Cam) with custom-made optical add-ons operated in event-mode data acquisition is…

cs.LG2022

Scaling Up Models and Data with and

Adam Roberts, Hyung Won Chung, Anselm Levskaya +40

Recent neural network-based language models have benefited greatly from scaling up the size of training datasets and the number of parameters in the models themselves. Scaling can…

physics.ins-det2021

A Novel Manufacturing Process for Glass THGEMs and First Characterisation in an Optical Gaseous Argon TPC

Adam Lowe, Krishanu Majumdar, Konstantinos Mavrokoridis +3

This paper details a novel, patent pending, abrasive machining manufacturing process for the formation of sub-millimetre holes in THGEMs, with the intended application in gaseous a…

cs.SD2018

Onsets and Frames: Dual-Objective Piano Transcription

Curtis Hawthorne, Erich Elsen, Jialin Song +6

We advance the state of the art in polyphonic piano music transcription by using a deep convolutional and recurrent neural network which is trained to jointly predict onsets and fr…

cs.CL2024

Training LLMs over Neurally Compressed Text

Brian Lester, Jaehoon Lee, Alex Alemi +4

In this paper, we explore the idea of training large language models (LLMs) over highly compressed text. While standard subword tokenizers compress text by a small factor, neural t…

cs.SD2025

Adaptive Accompaniment with ReaLchords

Yusong Wu, Tim Cooijmans, Kyle Kastner +10

Jamming requires coordination, anticipation, and collaborative creativity between musicians. Current generative models of music produce expressive output but are not able to genera…

cs.CL2023

Character-Aware Models Improve Visual Text Rendering

Rosanne Liu, Dan Garrette, Chitwan Saharia +7

Current image generation models struggle to reliably produce well-formed visual text. In this paper, we investigate a key contributing factor: popular text-to-image models lack cha…

cs.SD2019

GANSynth: Adversarial Neural Audio Synthesis

Jesse Engel, Kumar Krishna Agrawal, Shuo Chen +3

Efficient audio synthesis is an inherently difficult machine learning task, as human perception is sensitive to both global structure and fine-scale waveform coherence. Autoregress…