3 papers
cs.LG2026
Rethinking Personalized Reward Modeling for LLMs under Preference Heterogeneity via Group-Debiased Federated Learning
Seongyoon Kim, Boryeong Cho, Jihwan Oh +2
Large language models are increasingly aligned to human preferences via reward modeling, but user preference data are sensitive and often cannot be centralized. Federated learning…
cs.CV2026
Stage-wise Attention-Guided Region Sequencing for Adversarial Attacks on Large Vision-Language Models
Jaehyun Kwak, Nam Cao, Boryeong Cho +3
Targeted adversarial attacks on Large Vision-Language Models (LVLMs) test whether small image perturbations can steer model responses toward attacker-specified content. Under the s…
cs.CV2026
UniSAFE: A Comprehensive Benchmark for Safety Evaluation of Unified Multimodal Models
Segyu Lee, Boryeong Cho, Hojung Jung +8
Unified Multimodal Models (UMMs) offer powerful cross-modality capabilities but introduce new safety risks not observed in single-task models. Despite their emergence, existing saf…