7 papers
Delightful Exploration
Ian Osband
Most exploration algorithms search broadly until uncertainty is resolved. When the action space is too large to resolve within budget, practitioners default to -greedy…
Delightful Distributed Policy Gradient
Ian Osband
Distributed reinforcement learning trains on data from stale, buggy, or mismatched actors, producing actions with high surprisal (negative log-probability) under the learner's poli…
Delightful Gradients Accelerate Corner Escape
Jincheng Mei, Ian Osband
Softmax policy gradient converges at , but its transient behavior near sub-optimal corners of the simplex can be exponentially slow. The bottleneck is self-trapping: negati…
Does This Gradient Spark Joy?
Ian Osband
Policy gradient computes a backward pass for every sample, even though the backward pass is expensive and most samples carry little learning value. The Delightful Policy Gradient (…
Delightful Policy Gradient
Ian Osband
Standard policy gradients weight each sampled action by advantage alone, regardless of how likely that action was under the current policy. This creates two pathologies: within a s…
Balancing Knowledge Delivery and Emotional Comfort in Healthcare Conversational Systems
Shang-Chi Tsai, Yun-Nung Chen
With the advancement of large language models, many dialogue systems are now capable of providing reasonable and informative responses to patients' medical conditions. However, whe…