◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Tong Zhang

4 papers hereh-index 5275 citations5 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author1
  • last author3

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.LG3
  • cs.CL1
same name
  • Tong Zhang — 31 papers, h 31
  • Tong Zhang — 15 papers, h 23
  • Tong Zhang — 14 papers
  • Tong Zhang — 14 papers, h 4
  • Tong Zhang — 13 papers, h 47
  • Tong Zhang — 13 papers

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

collaborators

4 papers

cs.LG2024

Entropy-Regularized Process Reward Model

Hanning Zhang, Pengcheng Wang, Shizhe Diao +6

Large language models (LLMs) have shown promise in performing complex multi-step reasoning, yet they continue to struggle with mathematical reasoning, often making systematic error…

cs.LG2024

On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization

Yong Lin, Skyler Seto, Maartje ter Hoeve +6

Reinforcement Learning from Human Feedback (RLHF) is an effective approach for aligning language models to human preferences. Central to RLHF is learning a reward function for scor…

cs.LG2024

Localize-and-Stitch: Efficient Model Merging via Sparse Task Arithmetic

Yifei He, Yuzheng Hu, Yong Lin +2

Model merging offers an effective strategy to combine the strengths of multiple finetuned models into a unified model that preserves the specialized capabilities of each. Existing…

cs.CL2024

Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

Rui Yang, Ruomeng Ding, Yong Lin +2

Reward models trained on human preference data have been proven to effectively align Large Language Models (LLMs) with human intent within the framework of reinforcement learning f…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.