8 citations · 8 across the 2 of their papers we have counts for
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Text
Ximing Lu, David Acuna, Jaehun Jung +12
Reinforcement Learning with Verifiable Rewards (RLVR) has become a cornerstone for unlocking complex reasoning in Large Language Models (LLMs). Yet, scaling up RL is bottlenecked b…
cs.AI2022★ 8 cited
K-level Reasoning for Zero-Shot Coordination in Hanabi
Brandon Cui, Hengyuan Hu, Luis Pineda +1
The standard problem setting in cooperative multi-agent settings is self-play (SP), where the goal is to train a team of agents that works well together. However, optimal SP polici…
cs.AI2022
Self-Explaining Deviations for Coordination
Hengyuan Hu, Samuel Sokota, David Wu +4
Fully cooperative, partially observable multi-agent problems are ubiquitous in the real world. In this paper, we focus on a specific subclass of coordination problems in which huma…