◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Roberto-Rafael Maura-Rivero

3 papers hereh-index 223 citations4 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author1

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.AI1
  • cs.LG1
  • cs.MA1

identity via Semantic Scholar / OpenAlex

collaborators

3 papers

cs.LG2025

Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models

Roberto-Rafael Maura-Rivero, Chirag Nagpal, Roma Patel +1

Current methods that train large language models (LLMs) with reinforcement learning feedback, often resort to averaging outputs of multiple rewards functions during training. This…

cs.AI2025

Jackpot! Alignment as a Maximal Lottery

Roberto-Rafael Maura-Rivero, Marc Lanctot, Francesco Visin +1

Reinforcement Learning from Human Feedback (RLHF), the standard for aligning Large Language Models (LLMs) with human values, is known to fail to satisfy properties that are intuiti…

cs.MA2024

Soft Condorcet Optimization for Ranking of General Agents

Marc Lanctot, Kate Larson, Michael Kaisers +7

Driving progress of AI models and agents requires comparing their performance on standardized benchmarks; for general agents, individual performances must be aggregated across a po…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.