◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Qi Yao

4 papers hereh-index 358 citations6 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author3

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.AI3
  • cs.LG1
same name
  • Qi Yao — 1 paper, h 2
  • Qi Yao — 1 paper, h 2
  • Qi Yao — 1 paper, h 1
  • Qi Yao — 1 paper, h 2
  • Qi Yao — 1 paper, h 1
  • Qi Yao — 1 paper, h 8

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

collaborators

4 papers

cs.AI2025

Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction Following

Kongcheng Zhang, Qi Yao, Shunyu Liu +7

Reinforcement Learning (RL) has shown promise for aligning Large Language Models (LLMs) to follow instructions with various constraints. Despite the encouraging results, RL improve…

cs.AI2025

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning

Kongcheng Zhang, Qi Yao, Shunyu Liu +5

Recent advances of Reinforcement Learning (RL) have highlighted its potential in complex reasoning tasks, yet effective training often relies on external supervision, which limits…

cs.LG2025

Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood

Qingmao Yao, Zhichao Lei, Tianyuan Chen +5

Offline Reinforcement Learning (RL) struggles with distributional shifts, leading to the Q-value overestimation for out-of-distribution (OOD) actions. Existing methods address th…

cs.AI2025

Reasoning with Reinforced Functional Token Tuning

Kongcheng Zhang, Qi Yao, Baisheng Lai +5

In this work, we propose Reinforced Functional Token Tuning (RFTT), a novel reinforced fine-tuning framework that empowers Large Language Models (LLMs) with self-play learn-to-reas…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.