Showing 2026Show all
3 papers · 1 filter
cs.LG2026
Achieving Optimal Static and Dynamic Regret Simultaneously in Bandits with Deterministic Losses
Jian Qian, Chen-Yu Wei
In adversarial multi-armed bandits, two performance measures are commonly used: static regret, which compares the learner to the best fixed arm, and dynamic regret, which compares…
cs.GT2026
Is Online Linear Optimization Sufficient for Strategic Robustness?
Yang Cai, Haipeng Luo, Chen-Yu Wei +1
We consider bidding in repeated Bayesian first-price auctions. Bidding algorithms that achieve optimal regret have been extensively studied, but their strategic robustness to the s…
cs.LG2026
Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
Haolin Liu, Dian Yu, Sidi Lu +6
Reinforcement learning (RL) has emerged as a powerful framework for improving the reasoning capabilities of large language models (LLMs). However, most existing RL approaches rely…