2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.CL2024
To Burst or Not to Burst: Generating and Quantifying Improbable Text
Kuleen Sasse, Samuel Barham, Efsun Sarioglu Kayi +1
While large language models (LLMs) are extremely capable at text generation, their outputs are still distinguishable from human-authored text. We explore this separation across man…
cs.LG2023★ 2 cited
Clipped-Objective Policy Gradients for Pessimistic Policy Optimization
Jared Markowitz, Edward W. Staley
To facilitate efficient learning, policy gradient approaches to deep reinforcement learning (RL) are typically paired with variance reduction measures and strategies for making lar…