18 citations · 26 across the 10 of their papers we have counts for
10 papers
Flow-DPO: Improving LLM Mathematical Reasoning through Online Multi-Agent Learning
Yihe Deng, Paul Mineiro
Mathematical reasoning is a crucial capability for Large Language Models (LLMs), yet generating detailed and accurate reasoning traces remains a significant challenge. This paper i…
Online Joint Fine-tuning of Multi-Agent Flows
Paul Mineiro
A Flow is a collection of component models ("Agents") which constructs the solution to a complex problem via iterative communication. Flows have emerged as state of the art archite…
Efficient Contextual Bandits with Uninformed Feedback Graphs
Mengxiao Zhang, Yuheng Zhang, Haipeng Luo +1
Bandits with feedback graphs are powerful online learning models that interpolate between the full information and classic bandit problems, capturing many real-life applications. A…
Time-uniform confidence bands for the CDF under nonstationarity
Paul Mineiro, Steven R. Howard
Estimation of the complete distribution of a random variable is a useful primitive for both manual and automated decision making. This problem has received extensive attention in t…
Infinite Action Contextual Bandits with Reusable Data Exhaust
Mark Rucker, Yinglun Zhu, Paul Mineiro
For infinite action contextual bandits, smoothed regret and reduction to regression results in state-of-the-art online performance with computational cost independent of the action…
Contextual Bandits with Smooth Regret: Efficient Learning in Continuous Action Spaces
Yinglun Zhu, Paul Mineiro
Designing efficient general-purpose contextual bandit algorithms that work with large -- or even continuous -- action spaces would facilitate application to important scenarios suc…