1 citations · 1 across the 4 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Diffusion Language Models are Provably Optimal Parallel Samplers
Haozhe Jiang, Nika Haghtalab, Lijie Chen
Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive models for faster inference via parallel token generation. We provide a rigorous foundati…
cs.LG2025
Understanding In-context Learning of Addition via Activation Subspaces
Xinyan Hu, Kayo Yin, Michael I. Jordan +2
To perform few-shot learning, language models extract signals from a few input-label pairs, aggregate them into a learned prediction rule, and apply this rule to new inputs. How is…
cs.LG2024★ 1 cited
Theoretical limitations of multi-layer Transformer
Lijie Chen, Binghui Peng, Hongxun Wu
Transformers, especially the decoder-only variants, are the backbone of most modern large language models; yet we do not have much understanding of their expressive power except fo…