◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Zhi-Peng Ma

3 papers hereh-index 222 citations6 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author3

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.CL2
  • cs.AI1

identity via Semantic Scholar / OpenAlex

activity
20242026
most citedGRAIT: Gradient-Driven Refusal-Aware Instruction Tuning for Effective Hallucination Mitigation

1 citations · 1 across the 2 of their papers we have counts for

collaborators

3 papers

cs.AI2026

PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning

Yipeng Shi, Zhipeng Ma, Yue Wang +4

In long-horizon LLM agent reinforcement learning, weak policies often repeat similar failures, producing uninformative rollout trajectories and limiting effective policy optimizati…

cs.CL2025★ 1 cited

GRAIT: Gradient-Driven Refusal-Aware Instruction Tuning for Effective Hallucination Mitigation

Runchuan Zhu, Zinco Jiang, Jiang Wu +6

Refusal-Aware Instruction Tuning (RAIT) aims to enhance Large Language Models (LLMs) by improving their ability to refuse responses to questions beyond their knowledge, thereby red…

cs.CL2024

Utilize the Flow before Stepping into the Same River Twice: Certainty Represented Knowledge Flow for Refusal-Aware Instruction Tuning

Runchuan Zhu, Zhipeng Ma, Jiang Wu +4

Refusal-Aware Instruction Tuning (RAIT) enables Large Language Models (LLMs) to refuse to answer unknown questions. By modifying responses of unknown questions in the training data…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.