3 papers
cs.LG2026
PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization
Boryeong Cho, Sumyeong Ahn, Se-Young Yun
Direct Preference Optimization (DPO) simplifies alignment through pairwise comparisons but assumes all observed preferences are reliable. Real data often violates this assumption,…
cs.LG2026
Rethinking Personalized Reward Modeling for LLMs under Preference Heterogeneity via Group-Debiased Federated Learning
Seongyoon Kim, Boryeong Cho, Jihwan Oh +2
Large language models are increasingly aligned to human preferences via reward modeling, but user preference data are sensitive and often cannot be centralized. Federated learning…
cs.CV2026
UniSAFE: A Comprehensive Benchmark for Safety Evaluation of Unified Multimodal Models
Segyu Lee, Boryeong Cho, Hojung Jung +8
Unified Multimodal Models (UMMs) offer powerful cross-modality capabilities but introduce new safety risks not observed in single-task models. Despite their emergence, existing saf…