25 citations · 40 across the 6 of their papers we have counts for
4 papers · 1 filter
Miles v0.1: Production-Level Post-Training
RadixArk, :, Tom Chen +11
We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the clean design of slime, Miles designs each stage of the reinforcement-lear…
RLHF and IIA: Perverse Incentives
Wanqiao Xu, Shi Dong, Xiuyuan Lu +3
Existing algorithms for reinforcement learning from human feedback (RLHF) can incentivize responses at odds with preferences because they are based on models that assume independen…
Simple Agent, Complex Environment: Efficient Reinforcement Learning with Agent States
Shi Dong, Benjamin Van Roy, Zhengyuan Zhou
We design a simple reinforcement learning (RL) agent that implements an optimistic version of -learning and establish through regret analysis that this agent can operate with so…
Comments on the Du-Kakade-Wang-Yang Lower Bounds
Benjamin Van Roy, Shi Dong
Du, Kakade, Wang, and Yang recently established intriguing lower bounds on sample complexity, which suggest that reinforcement learning with a misspecified representation is intrac…