4 papers · 1 filter
An empirical study on the limitation of Transformers in program trace generation
Simeng Sun
We study Transformers on the task \emph{program trace generation} (PTG), where models produce step-by-step execution traces for synthetic programs. Unlike existing algorithmic prob…
SWAN-GPT: An Efficient and Scalable Approach for Long-Context Language Modeling
Krishna C. Puvvada, Faisal Ladhak, Santiago Akle Serrano +8
We present a decoder-only Transformer architecture that robustly generalizes to sequence lengths substantially longer than those seen during training. Our model, SWAN-GPT, interlea…
How much do contextualized representations encode long-range context?
Simeng Sun, Cheng-Ping Hsieh
We analyze contextual representations in neural autoregressive language models, emphasizing long-range contexts that span several thousand tokens. Our methodology employs a perturb…
Suri: Multi-constraint Instruction Following for Long-form Text Generation
Chau Minh Pham, Simeng Sun, Mohit Iyyer
Existing research on instruction following largely focuses on tasks with simple instructions and short responses. In this work, we explore multi-constraint instruction following fo…