4 papers
Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions
Matthew Riemer, Tommaso Tosato, Amin Memarian +4
This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavior and should be required to obtain behavioral cer…
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
Istabrak Abbes, Gopeshh Subbaraj, Matthew Riemer +6
Training large language models (LLMs) typically involves pre-training on massive corpora, only to restart the process entirely when new data becomes available. A more efficient and…
Handling Delay in Real-Time Reinforcement Learning
Ivan Anokhin, Rishav Rishav, Matthew Riemer +3
Real-time reinforcement learning (RL) introduces several challenges. First, policies are constrained to a fixed number of actions per second due to hardware limitations. Second, th…
Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference
Matthew Riemer, Gopeshh Subbaraj, Glen Berseth +1
Realtime environments change even as agents perform action inference and learning, thus requiring high interaction frequencies to effectively minimize regret. However, recent advan…