3 papers
cs.LO2026
The Network Structure of Mathlib
Xinze Li, Nanyun Peng, Simone Severini +1
The ongoing development of Lean 4's Mathlib has produced a macroscopic structural complexity that interweaves logical, mathematical, and infrastructural dependencies. We present a…
cs.LG2025
Convergence Theorems for Entropy-Regularized and Distributional Reinforcement Learning
Yash Jhaveri, Harley Wiltzer, Patrick Shafto +2
In the pursuit of finding an optimal policy, reinforcement learning (RL) methods generally ignore the properties of learned policies apart from their expected return. Thus, even wh…
cs.LG2024
Action Gaps and Advantages in Continuous-Time Distributional Reinforcement Learning
Harley Wiltzer, Marc G. Bellemare, David Meger +2
When decisions are made at high frequency, traditional reinforcement learning (RL) methods struggle to accurately estimate action values. In turn, their performance is inconsistent…