1 paper · 1 filter
Parker Whitfill, Stewy Slocum
Alignment techniques for LLMs rely on optimizing preference-based objectives -- where these preferences are typically elicited as ordinal, binary choices between responses. Recent…