2 papers
cs.CL2026
Parallel Token Prediction for Language Models
Felix Draxler, Justus Will, Farrin Marouf Sofian +3
Autoregressive decoding in language models is inherently slow, generating only one token per forward pass. We propose Parallel Token Prediction (PTP), a general-purpose framework f…
cs.LG2024
Progressive Compression with Universally Quantized Diffusion Models
Yibo Yang, Justus C. Will, Stephan Mandt
Diffusion probabilistic models have achieved mainstream success in many generative modeling tasks, from image generation to inverse problem solving. A distinct feature of these mod…