3 citations · 5 across the 7 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026★ 1 cited
SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
Geon-Hyeong Kim, Yu Jin Kim, Byoungjip Kim +4
As Large Language Models (LLMs) are increasingly deployed in real-world applications, balancing helpfulness and safety has become a central challenge. A natural approach is to inco…
cs.LG2025★ 2 cited
Process Reward Models That Think
Muhammad Khalifa, Rishabh Agarwal, Lajanugen Logeswaran +5
Step-by-step verifiers -- also known as process reward models (PRMs) -- are a key ingredient for test-time scaling. PRMs require step-level supervision, making them expensive to tr…