4 papers
Multiplayer Nash Preference Optimization
Fang Wu, Xu Huang, Weihao Xuan +8
Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human preferences. However, reward-based methods grou…
Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards
Fang Wu, Aaron Tu, Weihao Xuan +21
Reinforcement learning with verifiable rewards (RLVR) is a practical, scalable way to improve large language models on math, code, and other structured tasks. However, we argue tha…
Exponentiable locales, revisited
Xu Huang
We give a moderately motivated exposition of exponentiable locales and the construction of exponentials in , without assuming prior knowledge of exponential topologic…
A Coherence Construction for the Propositional Universe
Xu Huang
We record a particularly simple construction on top of Lumsdaine's local universes that allows for a Coquand-style universe of propositions with propositional extensionality to be…