2 papers
cs.LG2026
Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels
Hua Qu, Yifan Li, Xiaodong Yuan
Direct Preference Optimization (DPO) has become an important method for aligning large language models (LLMs) with human preferences because it removes the need for explicit reward…
cs.CL2024
Think-then-Act: A Dual-Angle Evaluated Retrieval-Augmented Generation
Yige Shen, Hao Jiang, Hua Qu +1
Despite their impressive capabilities, large language models (LLMs) often face challenges such as temporal misalignment and generating hallucinatory content. Enhancing LLMs with re…