12 citations · 35 across the 23 of their papers we have counts for
23 papers
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
Siyu Chen, Heejune Sheen, Tianhao Wang +1
In-context learning (ICL) is a cornerstone of large language model (LLM) functionality, yet its theoretical foundations remain elusive due to the complexity of transformer architec…
Unveiling the Statistical Foundations of Chain-of-Thought Prompting Methods
Xinyang Hu, Fengzhuo Zhang, Siyu Chen +1
Chain-of-Thought (CoT) prompting and its variants have gained popularity as effective methods for solving multi-step reasoning problems using pretrained large language models (LLMs…
Provable Statistical Rates for Consistency Diffusion Models
Zehao Dou, Minshuo Chen, Mengdi Wang +1
Diffusion models have revolutionized various application domains, including computer vision and audio generation. Despite the state-of-the-art performance, diffusion models are kno…
Pessimistic Value Iteration for Multi-Task Data Sharing in Offline Reinforcement Learning
Chenjia Bai, Lingxiao Wang, Jianye Hao +4
Offline Reinforcement Learning (RL) has shown promising results in learning a task-specific policy from a fixed dataset. However, successful offline RL often relies heavily on the…
Sample-efficient Learning of Infinite-horizon Average-reward MDPs with General Function Approximation
Jianliang He, Han Zhong, Zhuoran Yang
We study infinite-horizon average-reward Markov decision processes (AMDPs) in the context of general function approximation. Specifically, we propose a novel algorithmic framework…
Unveil Conditional Diffusion Models with Classifier-free Guidance: A Sharp Statistical Theory
Hengyu Fu, Zhuoran Yang, Mengdi Wang +1
Conditional diffusion models serve as the foundation of modern image synthesis and find extensive application in fields like computational biology and reinforcement learning. In th…