activity
20242026
most citedBenefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective

1 citations · 1 across the 3 of their papers we have counts for

collaborators

7 papers

cs.CL2026

Peer-Predictive Self-Training for Language Model Reasoning

Shi Feng, Hanlin Zhang, Fan Nie +2

Mechanisms for continued self-improvement of language models without external supervision remain an open challenge. We propose Peer-Predictive Self-Training (PST), a label-free fin…

cs.GT2026

Contextual Procurement Auctions with Bandit Learning

Yiling Chen, Shi Feng, Sadie Zhao

We study repeated procurement auctions in which producers have private costs and the platform must learn the context-dependent value of selecting each producer. We evaluate perform…

cs.AI20261 cited

Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective

Siwei Wang, Yifei Shen, Haoran Sun +5

Recent reinforcement learning (RL) methods have substantially enhanced the planning capabilities of Large Language Models (LLMs), yet the theoretical basis for their effectiveness…

cs.LG2025

Data Reliability Scoring

Yiling Chen, Shi Feng, Paul Kattuman +1

How can we assess the reliability of a dataset without access to ground truth? We introduce the problem of reliability scoring for datasets collected from potentially strategic sou…

cs.LG2024

ALPINE: Unveiling the Planning Capability of Autoregressive Learning in Language Models

Siwei Wang, Yifei Shen, Shi Feng +3

Planning is a crucial element of both human intelligence and contemporary large language models (LLMs). In this paper, we initiate a theoretical investigation into the emergence of…

cs.GT2024

Carrot and Stick: Eliciting Comparison Data and Beyond

Yiling Chen, Shi Feng, Fang-Yi Yu

Comparison data elicited from people are fundamental to many machine learning tasks, including reinforcement learning from human feedback for large language models and estimating r…