2 papers
cs.AI2025
Understanding and Enhancing the Planning Capability of Language Models via Multi-Token Prediction
Qimin Zhong, Hao Liao, Siwei Wang +4
Large Language Models (LLMs) have achieved impressive performance across diverse tasks but continue to struggle with learning transitive relations, a cornerstone for complex planni…
cs.LG2025
Offline Learning for Combinatorial Multi-armed Bandits
Xutong Liu, Xiangxiang Dai, Jinhang Zuo +4
The combinatorial multi-armed bandit (CMAB) is a fundamental sequential decision-making framework, extensively studied over the past decade. However, existing work primarily focuse…