4 citations · 11 across the 11 of their papers we have counts for
1 paper · 2 filters
Jiyoung Lee, Seungho Kim, Seunghyun Won +6
AI alignment refers to models acting towards human-intended goals, preferences, or ethical principles. Given that most large-scale deep learning models act as black boxes and canno…