1 paper
Prasad Mahadik, Adrians Skapars
It is becoming increasingly necessary to have monitors check for harmful behaviors during language model interactions, but text-only monitoring has not been sufficient. This is bec…