3 papers
cs.CL2026
Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale
Ang Li, Ben Liu, Bin Han +215
Efficient and scalable agentic intelligence requires models that can deliver both low-latency responses and strong reasoning capabilities while remaining practical to train, serve,…
cs.LG2025
No-Regret Linear Bandits under Gap-Adjusted Misspecification
Chong Liu, Dan Qiao, Ming Yin +2
This work studies linear bandits under a new notion of gap-adjusted misspecification and is an extension of Liu et al. (2023). When the underlying reward function is not linear, ex…
cs.LG2025
On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures
Ming Yin, Mengdi Wang, Yu-Xiang Wang
This article reviews the recent advances on the statistical foundation of reinforcement learning (RL) in the offline and low-adaptive settings. We will start by arguing why offline…