4 papers
MedLoc-R1: Performance-Aware Curriculum Reward Scheduling for GRPO-Based Medical Visual Grounding
Guangjing Yang, Ziyuan Qin, Chaoran Zhang +6
Medical visual grounding serves as a crucial foundation for fine-grained multimodal reasoning and interpretable clinical decision support. Despite recent advances in reinforcement…
Bitboard version of Tetris AI
Xingguo Chen, Pingshou Xiong, Zhenyu Luo +6
The efficiency of game engines and policy optimization algorithms is crucial for training reinforcement learning (RL) agents in complex sequential decision-making tasks, such as Te…
MindSpeed RL: Distributed Dataflow for Scalable and Efficient RL Training on Ascend NPU Cluster
Laingjun Feng, Chenyi Pan, Xinjie Guo +11
Reinforcement learning (RL) is a paradigm increasingly used to align large language models. Popular RL algorithms utilize multiple workers and can be modeled as a graph, where each…
AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training
Zhenyu Han, Ansheng You, Haibo Wang +16
Reinforcement learning (RL) has become a pivotal technology in the post-training phase of large language models (LLMs). Traditional task-colocated RL frameworks suffer from signifi…