6 papers
Prediction Under Imperfect Compression: A Theory of Approximate MDL
Qian Li, Xinyu Mao, Shang-Hua Teng +1
Minimum Description Length (MDL) formalizes the principle of Occam's razor by optimizing the total description length: . Fo…
Rethinking the Role of Positional Encoding: Sliding-Window Transformers without PE Remain Turing Complete
Qian Li, Xinyu Mao, Shang-Hua Teng
Positional encoding (PE) is widely viewed as necessary for transformers to process ordered sequences: without them, the next-token map appears permutation-invariant in its context…
A Theory of Time-Sensitive Language Generation: Sparse Hallucination Beats Mode Collapse
Atul Ganju, Travis McVoy, Shaddin Dughmi +1
We study language generation in the limit under a global preference ordering on strings, as introduced by Kleinberg and Wei. As is done in previous work, we aim for breadth, but im…
Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective
Siwei Wang, Yifei Shen, Haoran Sun +5
Recent reinforcement learning (RL) methods have substantially enhanced the planning capabilities of Large Language Models (LLMs), yet the theoretical basis for their effectiveness…
Proper Learnability and the Role of Unlabeled Data
Julian Asilis, Siddartha Devic, Shaddin Dughmi +2
Proper learning refers to the setting in which learners must emit predictors in the underlying hypothesis class , and often leads to learners with simple algorithmic forms (e.g.…
Semi-Random Graphs, Robust Asymmetry, and Reconstruction
Julian Asilis, Xi Chen, Dutch Hansen +1
The Graph Reconstruction Conjecture famously posits that any undirected graph on at least three vertices is determined up to isomorphism by its family of (unlabeled) induced subgra…