◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Bo Wang

7 papers hereh-index 5104 citations12 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author6

Across the 7 of 7 papers where every author was matched, so the position is known.

fields
  • cs.LG4
  • cs.CL2
  • cs.AI1
same name
  • Bo Wang — 75 papers, h 3
  • Bo Wang — 58 papers, h 5
  • Bo Wang — 32 papers
  • Bo Wang — 29 papers, h 22
  • Bo Wang — 17 papers, h 16
  • Bo Wang — 11 papers

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20242026
most citedScaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective

4 citations · 4 across the 4 of their papers we have counts for

collaborators
Showing cs.LGShow all

4 papers · 1 filter

cs.LG2026

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping

Yang Liu, Enxi Wang, Yufei Gao +6

Despite the success of reinforcement learning for large language models, a common failure mode is reduced sampling diversity, where the policy repeatedly generates similar erroneou…

cs.LG2026

BandPO: Bridging Trust Regions and Ratio Clipping via Probability-Aware Bounds for LLM Reinforcement Learning

Yuan Li, Bo Wang, Yufei Gao +4

Proximal constraints are fundamental to the stability of the Large Language Model reinforcement learning. While the canonical clipping mechanism in PPO serves as an efficient surro…

cs.LG2026

Explicit Multi-head Attention for Inter-head Interaction in Large Language Models

Runyu Peng, Yunhua Zhou, Demin Song +4

In large language models built upon the Transformer architecture, recent studies have shown that inter-head interaction can enhance attention performance. Motivated by this, we pro…

cs.LG2025

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections

Bo Wang, Qinyuan Cheng, Runyu Peng +7

Post-training processes are essential phases in grounding pre-trained language models to real-world tasks, with learning from demonstrations or preference signals playing a crucial…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.