Showing stat.MLShow all
2 papers · 1 filter
stat.ML2025
On Next-Token Prediction in LLMs: How End Goals Determine the Consistency of Decoding Algorithms
Jacob Trauger, Ambuj Tewari
Probabilistic next-token prediction trained using cross-entropy loss is the basis of most large language models. Given a sequence of previous values, next-token prediction assigns…
stat.ML2023
Sequence Length Independent Norm-Based Generalization Bounds for Transformers
Jacob Trauger, Ambuj Tewari
This paper provides norm-based generalization bounds for the Transformer architecture that do not depend on the input sequence length. We employ a covering number based approach to…