◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Jian Xie

5 papers hereh-index 4138 citations6 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author4
  • last author1

Across the 5 of 5 papers where every author was matched, so the position is known.

fields
  • cs.LG3
  • cs.AI1
  • cs.CL1
same name
  • Jian Xie — 16 papers, h 15
  • Jian Xie — 7 papers, h 3
  • Jian Xie — 5 papers, h 4
  • Jian Xie — 5 papers, h 5
  • Jian Xie — 3 papers, h 2
  • Jian Xie — 2 papers, h 1

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

collaborators

5 papers

cs.LG2025

Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown

Xingzhou Lou, Dong Yan, Wei Shen +3

Reward models (RMs) are essential for aligning large language models (LLM) with human expectations. However, existing RMs struggle to capture the stochastic and uncertain nature of…

cs.AI2025

3D-Properties: Identifying Challenges in DPO and Charting a Path Forward

Yuzi Yan, Yibo Miao, Jialian Li +4

Aligning large language models (LLMs) with human preferences has gained significant attention, with Proximal Policy Optimization (PPO) as a standard yet computationally expensive m…

cs.CL2025

Baichuan4-Finance Technical Report

Hanyu Zhang, Boyu Qiu, Yuhao Feng +6

Large language models (LLMs) have demonstrated strong capabilities in language understanding, generation, and reasoning, yet their potential in finance remains underexplored due to…

cs.LG2024

Boosting Deductive Reasoning with Step Signals In RLHF

Jialian Li, Yipin Zhang, Wei Shen +3

Logical reasoning is a crucial task for Large Language Models (LLMs), enabling them to tackle complex problems. Among reasoning tasks, multi-step reasoning poses a particular chall…

cs.LG2024

Reward-Robust RLHF in LLMs

Yuzi Yan, Xingzhou Lou, Jialian Li +6

As Large Language Models (LLMs) continue to progress toward more advanced forms of intelligence, Reinforcement Learning from Human Feedback (RLHF) is increasingly seen as a key pat…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.