Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Abstractive Red-Teaming of Language Model Character
Nate Rahn, Allison Qi, Avery Griffin +3
We want language model assistants to conform to a character specification, which asserts how the model should act across diverse user interactions. While models typically follow th…
cs.LG2023
Policy Optimization in a Noisy Neighborhood: On Return Landscapes in Continuous Control
Nate Rahn, Pierluca D'Oro, Harley Wiltzer +2
Deep reinforcement learning agents for continuous control are known to exhibit significant instability in their performance over time. In this work, we provide a fresh perspective…