◍wovepaper
SearchResearchersInstitutions
Sign in
institution

Huahai Pharmaceutical (China)

China

1 paper here
fields
  • cs.LG1
ROR 02jhay860OpenAlex

affiliations via OpenAlex

researchers with a paper here
  • Yu Wang1 · h 4

1 paper

cs.LG2026

The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and The Channel, Not the Content, Decides What Works

Yu Wang

Dense per-step supervision is the standard remedy for sparse-reward long-horizon LLM agents: reward the policy for predicting its next observation, which looks provably safe under…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.