3 papers
stat.ML2026
Policy Optimization and Statistical Inference for Online Contextual Matrix Games
Liner Xiang, Yixin Wang, Hengrui Cai
Online decision making often requires navigating a landscape shaped by both dynamic contexts and strategic interactions. In competitive pricing, for example, hotels must account fo…
cs.AI2026
Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge
Wenbo Zhang, Lijinghua Zhang, Liner Xiang +1
Reasoning-capable large language models (LLMs) have recently been adopted as automated judges, but their benefits and costs in LLM-as-a-Judge settings remain unclear. Through contr…
stat.ML2025
Foresighted Online Policy Optimization with Interference
Liner Xiang, Jiayi Wang, Hengrui Cai
Contextual bandits, which leverage the baseline features of sequentially arriving individuals to optimize cumulative rewards while balancing exploration and exploitation, are criti…