4 papers
Peer-Predictive Self-Training for Language Model Reasoning
Shi Feng, Hanlin Zhang, Fan Nie +2
Mechanisms for continued self-improvement of language models without external supervision remain an open challenge. We propose Peer-Predictive Self-Training (PST), a label-free fin…
Contextual Procurement Auctions with Bandit Learning
Yiling Chen, Shi Feng, Sadie Zhao
We study repeated procurement auctions in which producers have private costs and the platform must learn the context-dependent value of selecting each producer. We evaluate perform…
Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective
Siwei Wang, Yifei Shen, Haoran Sun +5
Recent reinforcement learning (RL) methods have substantially enhanced the planning capabilities of Large Language Models (LLMs), yet the theoretical basis for their effectiveness…
Data Reliability Scoring
Yiling Chen, Shi Feng, Paul Kattuman +1
How can we assess the reliability of a dataset without access to ground truth? We introduce the problem of reliability scoring for datasets collected from potentially strategic sou…