most citedDROP: Distributional and Regular Optimism and Pessimism for Reinforcement Learning

1 citations · 1 across the 2 of their papers we have counts for

collaborators

5 papers

cs.LG20261 cited

DROP: Distributional and Regular Optimism and Pessimism for Reinforcement Learning

Taisuke Kobayashi

In reinforcement learning (RL), temporal difference (TD) error is known to be related to the firing rate of dopamine neurons. It has been observed that each dopamine neuron does no…

cs.LG2026

Pseudo-Quantized Actor-Critic Algorithm for Robustness to Noisy Temporal Difference Error

Taisuke Kobayashi

In reinforcement learning (RL), temporal difference (TD) errors are widely adopted for optimizing value and policy functions. However, since the TD error is defined by a bootstrap…

cs.LG2025

Variational Adaptive Noise and Dropout towards Stable Recurrent Neural Networks

Taisuke Kobayashi, Shingo Murata

This paper proposes a novel stable learning theory for recurrent neural networks (RNNs), so-called variational adaptive noise and dropout (VAND). As stabilizing factors for RNNs, n…

cs.RO2025

LiRA: Light-Robust Adversary for Model-based Reinforcement Learning in Real World

Taisuke Kobayashi

Model-based reinforcement learning has attracted much attention due to its high sample efficiency and is expected to be applied to real-world robotic applications. In the real worl…

cs.LG2025

Improvements of Dark Experience Replay and Reservoir Sampling towards Better Balance between Consolidation and Plasticity

Taisuke Kobayashi

Continual learning is the one of the most essential abilities for autonomous agents, which can incrementally learn daily-life skills. For this ultimate goal, a simple but powerful…