most citedPhysicsFC: Learning User-Controlled Skills for a Physics-Based Football Player Controller

4 citations · 4 across the 2 of their papers we have counts for

collaborators

8 papers

cs.AI2026

GFlowPO: Generative Flow Network as a Language Model Prompt Optimizer

Junmo Cho, Suhan Kim, Sangjune An +5

Finding effective prompts for language models (LMs) is critical yet notoriously difficult: the prompt space is combinatorially large, rewards are sparse due to expensive target-LM…

cs.GR20254 cited

PhysicsFC: Learning User-Controlled Skills for a Physics-Based Football Player Controller

Minsu Kim, Eunho Jung, Yoonsang Lee

We propose PhysicsFC, a method for controlling physically simulated football player characters to perform a variety of football skills--such as dribbling, trapping, moving, and kic…

cs.AI2025

Self-Evolving Curriculum for LLM Reasoning

Xiaoyin Chen, Jiarui Lu, Minsu Kim +6

Reinforcement learning (RL) has proven effective for fine-tuning large language models (LLMs), significantly enhancing their reasoning abilities in domains such as mathematics and…

cs.LG2025

Adaptive Inference-Time Scaling via Cyclic Diffusion Search

Gyubin Lee, Truong Nhat Nguyen Bao, Jaesik Yoon +4

Diffusion models have demonstrated strong generative capabilities across domains ranging from image synthesis to complex reasoning tasks. However, most inference-time scaling metho…

cs.LG2025

Solving Bayesian inverse problems with diffusion priors and off-policy RL

Luca Scimeca, Siddarth Venkatraman, Moksh Jain +14

This paper presents a practical application of Relative Trajectory Balance (RTB), a recently introduced off-policy reinforcement learning (RL) objective that can asymptotically sol…

cs.LG2025

Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training

Brian Bartoldson, Siddarth Venkatraman, James Diffenderfer +7

Reinforcement learning (RL) is a critical component of large language model (LLM) post-training. However, on-policy algorithms used for post-training are not naturally robust to a…