activity
20242026
collaborators

7 papers

cs.CV2026

COMBAT: Conditional World Models for Behavioral Agent Training

Anmol Agarwal, Pranay Meshram, Sumer Singh +5

Recent advances in video generation have spurred the development of world models capable of simulating 3D-consistent environments and interactions with static objects. However, a s…

cs.LG2025

Results of the NeurIPS 2023 Neural MMO Competition on Multi-task Reinforcement Learning

Joseph Suárez, Kyoung Whan Choe, David Bloomin +22

We present the results of the NeurIPS 2023 Neural MMO Competition, which attracted over 200 participants and submissions. Participants trained goal-conditional policies that genera…

cs.LG2025

Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Alon Albalak, Duy Phung, Nathan Lile +8

Increasing interest in reasoning models has led math to become a prominent testing ground for algorithmic and methodological improvements. However, existing open math datasets eith…

cs.AI2025

Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought

Violet Xiang, Charlie Snell, Kanishk Gandhi +11

We propose a novel framework, Meta Chain-of-Thought (Meta-CoT), which extends traditional Chain-of-Thought (CoT) by explicitly modeling the underlying reasoning required to arrive…

cs.LG2024

Generative Reward Models

Dakota Mahan, Duy Van Phung, Rafael Rafailov +6

Reinforcement Learning from Human Feedback (RLHF) has greatly improved the performance of modern Large Language Models (LLMs). The RLHF process is resource-intensive and technicall…

cs.CL2024

Self-Directed Synthetic Dialogues and Revisions Technical Report

Nathan Lambert, Hailey Schoelkopf, Aaron Gokaslan +3

Synthetic data has become an important tool in the fine-tuning of language models to follow instructions and solve complex problems. Nevertheless, the majority of open data to date…