activity
20242026
collaborators

10 papers

hep-ex2026

Predict before you train: Scaling Laws for particle physics foundation models

Jan-Lucas Uslu, Benjamin Nachman, Christopher Re

The largest machine learning models in particle physics are also the most expensive to train, yet the return on scaling a given architecture cannot be estimated before that compute…

cs.CV2025

HMAR: Efficient Hierarchical Masked Auto-Regressive Image Generation

Hermann Kumbong, Xian Liu, Tsung-Yi Lin +6

Visual Auto-Regressive modeling (VAR) has shown promise in bridging the speed and quality gap between autoregressive image models and diffusion models. VAR reformulates autoregress…

cs.LG2025

Restructuring Vector Quantization with the Rotation Trick

Christopher Fifty, Ronald G. Junkins, Dennis Duan +5

Vector Quantized Variational AutoEncoders (VQ-VAEs) are designed to compress a continuous input to a discrete latent space and reconstruct it with minimal distortion. They operate…

cs.LG2025

Towards Learning High-Precision Least Squares Algorithms with Sequence Models

Jerry Liu, Jessica Grogan, Owen Dugan +4

This paper investigates whether sequence models can learn to perform numerical algorithms, e.g. gradient descent, on the fundamental problem of least squares. Our goal is to inheri…

cs.LG2025

LoLCATs: On Low-Rank Linearizing of Large Language Models

Michael Zhang, Simran Arora, Rahul Chalamala +5

Recent works show we can linearize large language models (LLMs) -- swapping the quadratic attentions of popular Transformer-based LLMs with subquadratic analogs, such as linear att…

cs.LG2025

Systems and Algorithms for Convolutional Multi-Hybrid Language Models at Scale

Jerome Ku, Eric Nguyen, David W. Romero +13

We introduce convolutional multi-hybrid architectures, with a design grounded on two simple observations. First, operators in hybrid models can be tailored to token manipulation ta…