Showing stat.MLShow all
2 papers · 1 filter
stat.ML2026
Tail-Aware Information-Theoretic Bounds for LLM Alignment under Heavy-Tailed Rewards
Huiming Zhang, Binghan Li, Wan Tian +1
Classical information-theoretic learning bounds typically rely on KL mutual information and moment-generating-function (MGF) arguments, which are well matched to bounded or sub-Gau…
stat.ML2024
Selective Reviews of Bandit Problems in AI via a Statistical View
Pengjie Zhou, Haoyu Wei, Huiming Zhang
Reinforcement Learning (RL) is a widely researched area in artificial intelligence that focuses on teaching agents decision-making through interactions with their environment. A ke…