◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Mitchell Keith Bloch

2 papers hereh-index 27 citations6 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • sole author2

Across the 2 of 2 papers where every author was matched, so the position is known.

fields
  • cs.LG2

identity via Semantic Scholar / OpenAlex

most citedReducing Commitment to Tasks with Off-Policy Hierarchical Reinforcement Learning

3 citations · 5 across the 2 of their papers we have counts for

collaborators

2 papers

cs.LG2011★ 3 cited

Reducing Commitment to Tasks with Off-Policy Hierarchical Reinforcement Learning

Mitchell Keith Bloch

In experimenting with off-policy temporal difference (TD) methods in hierarchical reinforcement learning (HRL) systems, we have observed unwanted on-policy learning under reproduci…

cs.LG2011★ 2 cited

Temporal Second Difference Traces

Mitchell Keith Bloch

Q-learning is a reliable but inefficient off-policy temporal-difference method, backing up reward only one step at a time. Replacing traces, using a recency heuristic, are more eff…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.