1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Shir Rozenfeld, Rahul Pankajakshan, Itay Zloczower +3
Large language models (LLMs) are increasingly paired with activation-based monitoring to detect and prevent harmful behaviors that may not be apparent at the surface-text level. Ho…