4 papers
Few Batches or Little Memory, But Not Both: Simultaneous Space and Adaptivity Constraints in Stochastic Bandits
Ruiyuan Huang, Zicheng Lyu, Xiaoyi Zhu +1
We study stochastic multi-armed bandits under simultaneous constraints on space and adaptivity: the learner interacts with the environment in batches and has only bits of p…
EvolProver: Advancing Automated Theorem Proving by Evolving Formalized Problems via Symmetry and Difficulty
Yuchen Tian, Ruiyuan Huang, Xuanwu Wang +6
Large Language Models (LLMs) for formal theorem proving have shown significant promise, yet they often lack generalizability and are fragile to even minor transformations of proble…
Nearly Tight Bounds for Cross-Learning Contextual Bandits with Graphical Feedback
Ruiyuan Huang, Zengfeng Huang
Repeated first-price auctions are contextual decision problems with censored but reusable feedback: after submitting a bid, a learner can infer the outcomes of related bids and eva…
High Probability Bound for Cross-Learning Contextual Bandits with Unknown Context Distributions
Ruiyuan Huang, Zengfeng Huang
Motivated by applications in online bidding and sleeping bandits, we examine the problem of contextual bandits with cross learning, where the learner observes the loss associated w…