◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Jimmy Ba

2 papers hereh-index 51.5k citations11 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • last author1

Across the 1 of 2 papers where every author was matched, so the position is known.

fields
  • cs.LG2
same name
  • Jimmy Ba — 2 papers, h 2
  • Jimmy Ba — 2 papers, h 46
  • Jimmy Ba — 1 paper, h 3

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

collaborators

2 papers

cs.LG2026

EMA Policy Gradient: Taming Reinforcement Learning for LLMs with EMA Anchor and Top-k KL

Lunjun Zhang, Jimmy Ba

Reinforcement Learning (RL) has enabled Large Language Models (LLMs) to acquire increasingly complex reasoning and agentic behaviors. In this work, we propose two simple techniques…

cs.LG2024

The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning

Nathaniel Li, Alexander Pan, Anjali Gopal +54

The White House Executive Order on Artificial Intelligence highlights the risks of large language models (LLMs) empowering malicious actors in developing biological, cyber, and che…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.