4 citations · 4 across the 6 of their papers we have counts for
11 papers
Scaling Domain Data Repetition in LLM Pretraining
Jingwei Li, Xinran Gu, Rui Dai +5
As large language models scale, their training-token budgets must also increase to maintain an appropriate tokens-per-parameter ratio (\(\mathrm{TPP}\)). However, high-quality doma…
Explaining Data Mixing Scaling Laws
Rui Dai, Shuran Zheng
Recent research has established empirical scaling laws to predict model performance on multi-domain data mixtures. However, a theoretical understanding of these model loss behavior…
Bayesian Conversations
Renato Paes Leme, Jon Schneider, Heyang Shang +1
We initiate the study of Bayesian conversations, which model interactive communication between two strategic agents without a mediator. We compare this to communication through a m…
Federated Learning as a Network Effects Game
Shengyuan Hu, Dung Daniel Ngo, Shuran Zheng +2
Federated Learning (FL) aims to foster collaboration among a population of clients to improve the accuracy of machine learning without directly sharing local data. Although there h…
Private Interdependent Valuations
Alon Eden, Kira Goldner, Shuran Zheng
We consider the single-item interdependent value setting, where there is a monopolist, buyers, and each buyer has a private signal describing a piece of information about…
The Limits of Multi-task Peer Prediction
Shuran Zheng, Fang-Yi Yu, Yiling Chen
Recent advances in multi-task peer prediction have greatly expanded our knowledge about the power of multi-task peer prediction mechanisms. Various mechanisms have been proposed in…