1 citations · 1 across the 17 of their papers we have counts for
1 paper · 1 filter
Jihyung Park, Saleh Afroogh, Junfeng Jiao
Current language models create two safety challenges: risk must be detected early enough to avoid exposing harmful continuation, and the harmfulness itself may be implicit rather t…