2 papers
cs.CL2026
Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference Optimization
Junming Yang, Ning Xu, Biao Liu +2
Preference optimization is crucial for aligning large language models (LLMs) with human values and intentions. A significant challenge in this process is the distribution mismatch…
cs.CL2024
Negative-Prompt-driven Alignment for Generative Language Model
Shiqi Qiao, Ning Xv, Biao Liu +1
Large language models have achieved remarkable capabilities, but aligning their outputs with human values and preferences remains a significant challenge. Existing alignment method…