2 papers
cs.AI2025
Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models
Xin Zhou, Yiwen Guo, Ruotian Ma +3
Aligning Large Language Models (LLMs) with human preferences is crucial for their deployment in real-world applications. Recent advancements in Self-Rewarding Language Models sugge…
cs.AI2025
Reasoning-as-Logic-Units: Scaling Test-Time Reasoning in Large Language Models Through Logic Unit Alignment
Cheryl Li, Tianyuan Xu, Yiwen Guo
Chain-of-Thought (CoT) prompting has shown promise in enhancing the reasoning capabilities of large language models (LLMs) by generating natural language (NL) rationales that lead…