◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Wenbo Su

4 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author4

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.LG2
  • cs.AI1
  • cs.IR1
same name
  • Wenbo Su — 15 papers
  • Wenbo Su — 3 papers
  • Wenbo Su — 1 paper
  • Wenbo Su — 1 paper

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

most citedEquilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models

1 citations · 1 across the 4 of their papers we have counts for

collaborators

4 papers

cs.LG2025

Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning

Jiashun Liu, Johan Obando-Ceron, Han Lu +7

Most recent RL for LLMs (RL4LLM) methods avoid explicit critics, replacing them with average advantage baselines. This shift is largely pragmatic: conventional value functions are…

cs.LG2025

Part II: ROLL Flash -- Accelerating RLVR and Agentic Training with Asynchrony

Han Lu, Zichen Liu, Shaopan Xiong +19

Synchronous Reinforcement Learning (RL) post-training has emerged as a crucial step for enhancing Large Language Models (LLMs) with diverse capabilities. However, many systems desi…

cs.IR2025

MIM: Multi-modal Content Interest Modeling Paradigm for User Behavior Modeling

Bencheng Yan, Si Chen, Shichang Jia +12

Click-Through Rate (CTR) prediction is a crucial task in recommendation systems, online searches, and advertising platforms, where accurately capturing users' real interests in con…

cs.AI2025★ 1 cited

Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models

Yingshui Tan, Yilei Jiang, Yanshi Li +6

Fine-tuning large language models (LLMs) based on human preferences, commonly achieved through reinforcement learning from human feedback (RLHF), has been effective in improving th…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.