2 papers
cs.CL2025
Atomic Consistency Preference Optimization for Long-Form Question Answering
Jingfeng Chen, Raghuveer Thirukovalluru, Junlin Wang +2
Large Language Models (LLMs) often produce factoid hallucinations - plausible yet incorrect answers. A common mitigation strategy is model alignment, which improves factual accurac…
cs.CL2025
Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards
Ruipeng Jia, Yunyi Yang, Yongbo Gai +5
Reinforcement learning with verifiable rewards (RLVR) has enabled large language models (LLMs) to achieve remarkable breakthroughs in reasoning tasks with objective ground-truth an…