4 papers
ARCA: Adapter-Residual Credit Assignment When Token Signals Degenerate
Rodney Lafuente-Mercado
Token-level credit assignment for language-model reinforcement learning is usually formulated as if the policy were fully trainable, while practical LLM-RL pipelines often rely on…
When Sensors Fail: Temporal Sequence Models for Robust PPO under Sensor Drift
Kevin Vogt-Lowell, Theodoros Tsiligkaridis, Rodney Lafuente-Mercado +4
Real-world reinforcement learning systems must operate under distributional drift in their observation streams, yet most policy architectures implicitly assume fully observed and n…
Adaptive Policy Synchronization for Scalable Reinforcement Learning
Rodney Lafuente-Mercado
Scaling reinforcement learning (RL) often requires running environments across many machines, but most frameworks tie simulation, training, and infrastructure into rigid systems. W…
Syndeo: Portable Ray Clusters with Secure Containerization
William Li, Rodney S. Lafuente Mercado, Jaime D. Pena +1
We present Syndeo: a software framework for container orchestration of Ray on Slurm. In general the idea behind Syndeo is to write code once and deploy anywhere. Specifically, Synd…