◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Zi-Qiang Dong

3 papers hereh-index 210 citations7 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author1
  • last author1

Across the 2 of 3 papers where every author was matched, so the position is known.

fields
  • cs.LG2
  • cs.AI1

identity via Semantic Scholar / OpenAlex

collaborators

3 papers

cs.LG2026

I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization

Yubo Zhang, Xinhong Ma, Zezhong Tan +1

Group Relative Policy Optimization (GRPO) learns from reward differences within a rollout group, but receives no useful relative signal when every sampled response is incorrect. Pr…

cs.LG2026

Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation

Zhilin Huang, Hang Gao, Ziqiang Dong +6

Self-distillation improves reasoning in large language models by using the model's own rollouts as training signal, typically through implicit logit-level alignment that minimizes…

cs.AI2025

Towards Flash Thinking via Decoupled Advantage Policy Optimization

Zezhong Tan, Hang Gao, Xinhong Ma +2

Recent Large Reasoning Models (LRMs) have achieved remarkable performance in solving complex problems via supervised fine-tuning (SFT) and reinforcement learning (RL). Although exi…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.