1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.LG2025
LLM Pretraining with Continuous Concepts
Jihoon Tack, Jack Lanchantin, Jane Yu +7
Next token prediction has been the standard training objective used in large language model pretraining. Representations are learned as a result of optimizing for token-level perpl…
cs.LG2024★ 1 cited
Offline Reinforcement Learning for LLM Multi-Step Reasoning
Huaijie Wang, Shibo Hao, Hanze Dong +4
Improving the multi-step reasoning ability of large language models (LLMs) with offline reinforcement learning (RL) is essential for quickly adapting them to complex tasks. While D…