activity
20142024
most citedA Randomized Algorithm for CCA

18 citations · 26 across the 10 of their papers we have counts for

collaborators

10 papers

cs.CL2024

Flow-DPO: Improving LLM Mathematical Reasoning through Online Multi-Agent Learning

Yihe Deng, Paul Mineiro

Mathematical reasoning is a crucial capability for Large Language Models (LLMs), yet generating detailed and accurate reasoning traces remains a significant challenge. This paper i…

cs.LG2024

Online Joint Fine-tuning of Multi-Agent Flows

Paul Mineiro

A Flow is a collection of component models ("Agents") which constructs the solution to a complex problem via iterative communication. Flows have emerged as state of the art archite…

cs.LG2024

Efficient Contextual Bandits with Uninformed Feedback Graphs

Mengxiao Zhang, Yuheng Zhang, Haipeng Luo +1

Bandits with feedback graphs are powerful online learning models that interpolate between the full information and classic bandit problems, capturing many real-life applications. A…

stat.ML2023

Time-uniform confidence bands for the CDF under nonstationarity

Paul Mineiro, Steven R. Howard

Estimation of the complete distribution of a random variable is a useful primitive for both manual and automated decision making. This problem has received extensive attention in t…

cs.LG2023

Infinite Action Contextual Bandits with Reusable Data Exhaust

Mark Rucker, Yinglun Zhu, Paul Mineiro

For infinite action contextual bandits, smoothed regret and reduction to regression results in state-of-the-art online performance with computational cost independent of the action…

cs.LG20222 cited

Contextual Bandits with Smooth Regret: Efficient Learning in Continuous Action Spaces

Yinglun Zhu, Paul Mineiro

Designing efficient general-purpose contextual bandit algorithms that work with large -- or even continuous -- action spaces would facilitate application to important scenarios suc…