4 citations · 4 across the 2 of their papers we have counts for
8 papers
GFlowPO: Generative Flow Network as a Language Model Prompt Optimizer
Junmo Cho, Suhan Kim, Sangjune An +5
Finding effective prompts for language models (LMs) is critical yet notoriously difficult: the prompt space is combinatorially large, rewards are sparse due to expensive target-LM…
PhysicsFC: Learning User-Controlled Skills for a Physics-Based Football Player Controller
Minsu Kim, Eunho Jung, Yoonsang Lee
We propose PhysicsFC, a method for controlling physically simulated football player characters to perform a variety of football skills--such as dribbling, trapping, moving, and kic…
Self-Evolving Curriculum for LLM Reasoning
Xiaoyin Chen, Jiarui Lu, Minsu Kim +6
Reinforcement learning (RL) has proven effective for fine-tuning large language models (LLMs), significantly enhancing their reasoning abilities in domains such as mathematics and…
Adaptive Inference-Time Scaling via Cyclic Diffusion Search
Gyubin Lee, Truong Nhat Nguyen Bao, Jaesik Yoon +4
Diffusion models have demonstrated strong generative capabilities across domains ranging from image synthesis to complex reasoning tasks. However, most inference-time scaling metho…
Solving Bayesian inverse problems with diffusion priors and off-policy RL
Luca Scimeca, Siddarth Venkatraman, Moksh Jain +14
This paper presents a practical application of Relative Trajectory Balance (RTB), a recently introduced off-policy reinforcement learning (RL) objective that can asymptotically sol…
Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training
Brian Bartoldson, Siddarth Venkatraman, James Diffenderfer +7
Reinforcement learning (RL) is a critical component of large language model (LLM) post-training. However, on-policy algorithms used for post-training are not naturally robust to a…