◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Satinder Singh

48 papers hereh-index 283.7k citations67 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author16
  • last author30

Across the 46 of 48 papers where every author was matched, so the position is known.

fields
  • cs.LG31
  • cs.AI11
  • cs.CL3
  • cs.NE1
  • cs.RO1
  • stat.ML1
same name
  • Satinder Singh — 15 papers, h 66
  • Satinder Singh — 7 papers, h 9
  • Satinder Singh — 4 papers, h 2
  • Satinder Singh — 3 papers, h 2
  • Satinder Singh — 2 papers
  • Satinder Singh — 2 papers, h 3

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20182023
most citedMastering the Game of Stratego with Model-Free Multiagent Reinforcement Learning

153 citations · 273 across the 25 of their papers we have counts for

collaborators
Showing 2021 · cs.AIShow all

4 papers · 2 filters

cs.AI2021

Proper Value Equivalence

Christopher Grimm, André Barreto, Gregory Farquhar +2

One of the main challenges in model-based reinforcement learning (RL) is to decide which aspects of the environment should be modeled. The value-equivalence (VE) principle proposes…

cs.AI2021

Discovering Diverse Nearly Optimal Policies with Successor Features

Tom Zahavy, Brendan O'Donoghue, Andre Barreto +3

Finding different solutions to the same problem is a key aspect of intelligence associated with creativity and adaptation to novel situations. In reinforcement learning, a set of d…

cs.AI2021

Reward is enough for convex MDPs

Tom Zahavy, Brendan O'Donoghue, Guillaume Desjardins +1

Maximising a cumulative reward function that is Markov and stationary, i.e., defined over state-action pairs and independent of time, is sufficient to capture many kinds of goals i…

cs.AI2021

Discovering a set of policies for the worst case reward

Tom Zahavy, Andre Barreto, Daniel J Mankowitz +4

We study the problem of how to construct a set of policies that can be composed together to solve a collection of reinforcement learning tasks. Each task is a different reward func…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.