1 paper · 1 filter
Arushi Rai, Qiang Zhang, Hanqing Zeng +5
Large language models (LLMs) exhibit strong reasoning capabilities but typically require expensive post-training to reach high performance. Recent test-time alignment methods offer…