2 papers
cs.LG2025
LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
Bo Wu, Sid Wang, Yunhao Tang +11
Reinforcement Learning (RL) has become the most effective post-training approach for improving the capabilities of Large Language Models (LLMs). In practice, because of the high de…
math.NA2023
Implementation and (Inverse Modified) Error Analysis for implicitly-templated ODE-nets
Aiqing Zhu, Tom Bertalan, Beibei Zhu +2
We focus on learning unknown dynamics from data using ODE-nets templated on implicit numerical initial value problem solvers. First, we perform Inverse Modified error analysis of t…