3 papers
cs.LG2026
A Relative-Budget Theory for Reinforcement Learning with Verifiable Rewards in Large Language Model Reasoning
Akifumi Wachi, Hirota Kinoshita, Shokichi Takakura +2
Reinforcement learning (RL) is a dominant paradigm for improving the reasoning abilities of large language models, yet its effectiveness varies across tasks and compute budgets. We…
cs.GT2024
Achieving PAC Guarantees in Mechanism Design through Multi-Armed Bandits
Takayuki Osogami, Hirota Kinoshita, Segev Wasserkrug
We analytically derive a class of optimal solutions to a linear program (LP) for automated mechanism design that satisfies efficiency, incentive compatibility, strong budget balanc…
cs.DS2024
A Faster Deterministic Algorithm for Mader's -Path Packing
Satoru Iwata, Hirota Kinoshita
Given an undirected graph with a set of terminals partitioned into a family of disjoint blocks, find the maximum number of vertex-disjoint…