◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Lingling Fu

2 papers hereh-index 18 citations9 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2

Across the 2 of 2 papers where every author was matched, so the position is known.

fields
  • cs.IR1
  • cs.LG1

identity via Semantic Scholar / OpenAlex

collaborators

2 papers

cs.IR2026

TPMM-DPO: Trajectory-aware Preference-guided Model Merging for Iterative Direct Preference Optimization

Lingling Fu, Yongfu Xu

Direct Preference Optimization (DPO) has been widely adopted for large language model alignment due to its simple training procedure and lack of an explicit reward model. However,…

cs.LG2026

UMM-RM: An Upcycle-and-Merge MoE Reward Model for Mitigating Reward Hacking

Lingling Fu, Yongfu Xue

Reward models (RMs) are a critical component of reinforcement learning from human feedback (RLHF). However, conventional dense RMs are susceptible to exploitation by policy models…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.