Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Magnetic Preference Optimization: Achieving Last-iterate Convergence for Language Model Alignment
Mingzhi Wang, Chengdong Ma, Qizhi Chen +7
Self-play methods have demonstrated remarkable success in enhancing model capabilities across various domains. In the context of Reinforcement Learning from Human Feedback (RLHF),…
cs.CL2024
J2N -- Nominal Adjective Identification and its Application
Lemeng Qi, Yang Han, Zhuotong Xie
This paper explores the challenges posed by nominal adjectives (NAs) in natural language processing (NLP) tasks, particularly in part-of-speech (POS) tagging. We propose treating N…