◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Wen Sun

5 papers hereh-index 583 citations8 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author1
  • last author4

Across the 5 of 5 papers where every author was matched, so the position is known.

fields
  • cs.LG3
  • cs.AI1
  • stat.ML1
same name
  • Wen Sun — 12 papers, h 6
  • Wen Sun — 11 papers, h 7
  • Wen Sun — 3 papers, h 2
  • Wen Sun — 3 papers, h 2
  • Wen Sun — 2 papers, h 0
  • Wen Sun — 2 papers, h 3

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

collaborators
Showing cs.LGShow all

4 papers · 1 filter

cs.LG2025

Q♯: Provably Optimal Distributional RL for LLM Post-Training

Jin Peng Zhou, Kaiwen Wang, Jonathan Chang +5

Reinforcement learning (RL) post-training is crucial for LLM alignment and reasoning, but existing policy-based methods, such as PPO and DPO, can fall short of fixing shortcuts inh…

cs.LG2025

Value-Guided Search for Efficient Chain-of-Thought Reasoning

Kaiwen Wang, Jin Peng Zhou, Jonathan Chang +4

In this paper, we propose a simple and efficient method for value model training on long-context reasoning traces. Compared to existing process reward models (PRMs), our method doe…

cs.LG2025

A Reductions Approach to Risk-Sensitive Reinforcement Learning with Optimized Certainty Equivalents

Kaiwen Wang, Dawen Liang, Nathan Kallus +1

We study risk-sensitive RL where the goal is learn a history-dependent policy that optimizes some risk measure of cumulative rewards. We consider a family of risks called the optim…

cs.LG2024

More Benefits of Being Distributional: Second-Order Bounds for Reinforcement Learning

Kaiwen Wang, Owen Oertell, Alekh Agarwal +2

In this paper, we prove that Distributional Reinforcement Learning (DistRL), which learns the return distribution, can obtain second-order bounds in both online and offline RL in g…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.