Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings
Yuning Wu, Ke Wang, Devin Chen +1
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for post-training reasoning models. However, group-based methods such as Group Relative Po…
cs.LG2024
Practical Marketplace Optimization at Uber Using Causally-Informed Machine Learning
Bobby Chen, Siyu Chen, Jason Dowlatabadi +15
Budget allocation of marketplace levers, such as incentives for drivers and promotions for riders, has long been a technical and business challenge at Uber; understanding lever bud…