5 papers
Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale
Ang Li, Ben Liu, Bin Han +215
Efficient and scalable agentic intelligence requires models that can deliver both low-latency responses and strong reasoning capabilities while remaining practical to train, serve,…
No-Regret Linear Bandits under Gap-Adjusted Misspecification
Chong Liu, Dan Qiao, Ming Yin +2
This work studies linear bandits under a new notion of gap-adjusted misspecification and is an extension of Liu et al. (2023). When the underlying reward function is not linear, ex…
On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures
Ming Yin, Mengdi Wang, Yu-Xiang Wang
This article reviews the recent advances on the statistical foundation of reinforcement learning (RL) in the offline and low-adaptive settings. We will start by arguing why offline…
Offline Multitask Representation Learning for Reinforcement Learning
Haque Ishfaq, Thanh Nguyen-Tang, Songtao Feng +4
We study offline multitask representation learning in reinforcement learning (RL), where a learner is provided with an offline dataset from different tasks that share a common repr…
NetworkGym: Reinforcement Learning Environments for Multi-Access Traffic Management in Network Simulation
Momin Haider, Ming Yin, Menglei Zhang +3
Mobile devices such as smartphones, laptops, and tablets can often connect to multiple access networks (e.g., Wi-Fi, LTE, and 5G) simultaneously. Recent advancements facilitate sea…