2 papers
cs.LG2026
Intrinsic Mutual Information as a Modulator for Preference Optimization
Peng Liao, Peijia Zheng, Lingbo Li +2
Offline preference optimization methods, such as Direct Preference Optimization (DPO), offer significant advantages in aligning Large Language Models (LLMs) with human values. Howe…
cs.AI2025
Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs
Mohammad Ali Alomrani, Yingxue Zhang, Derek Li +14
Large language models (LLMs) have rapidly progressed into general-purpose agents capable of solving a broad spectrum of tasks. However, current models remain inefficient at reasoni…