◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Ziyi Yang

4 papers hereh-index 4107 citations6 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author2

Across the 3 of 4 papers where every author was matched, so the position is known.

fields
  • cs.CL2
  • cs.LG2
same name
  • Ziyi Yang — 12 papers, h 2
  • Ziyi Yang — 10 papers, h 5
  • Ziyi Yang — 8 papers, h 13
  • Ziyi Yang — 7 papers
  • Ziyi Yang — 6 papers, h 1
  • Ziyi Yang — 6 papers, h 4

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

most citedPhi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle

2 citations · 2 across the 3 of their papers we have counts for

collaborators

4 papers

cs.CL2026

GGC: Selective Query Correction for Reliable Text-to-SPARQL Generation

Ziyi Yang, Thanh-Son Nguyen, Tuan Anh Nguyen +1

Large language models (LLMs) have demonstrated strong capabilities in structured query generation, making them a natural choice for Text-to-SPARQL, which translates natural languag…

cs.CL2024★ 2 cited

Phi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle

Emman Haider, Daniel Perez-Becker, Thomas Portet +28

Recent innovations in language model training have demonstrated that it is possible to create highly performant models that are small enough to run on a smartphone. As these models…

cs.LG2024

Cost-Effective Proxy Reward Model Construction with On-Policy and Active Learning

Yifang Chen, Shuohang Wang, Ziyi Yang +6

Reinforcement learning with human feedback (RLHF), as a widely adopted approach in current large language model pipelines, is \textit{bottlenecked by the size of human preference d…

cs.LG2024

Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Shenao Zhang, Donghan Yu, Hiteshi Sharma +6

Preference optimization, particularly through Reinforcement Learning from Human Feedback (RLHF), has achieved significant success in aligning Large Language Models (LLMs) to adhere…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.