◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Shao-Hua Ma

3 papers hereh-index 00 citations3 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • last author3

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.LG2
  • cs.AI1

identity via Semantic Scholar / OpenAlex

collaborators

3 papers

cs.LG2026

VICT: Verifier-Instrumented Credit Tracing for Long-Horizon LLM Agent Reinforcement Learning

Pengcheng Li, Zhengyang Zhang, Dongxu Zhang +2

Fine-grained credit assignment is a central challenge in reinforcement learning for long horizon LLM agents. Standard objectives often train from programmatically verifiable termin…

cs.AI2026

CARGO-VL: Counterfactual Arbitration with Risk-Constrained Group Optimization for Vision-Language Models

De Jiang, Zhengyang Zhang, Kehong Yuan +1

Vision-language systems combine images with retrieved text, but these sources can disagree or jointly fail to support an answer. Reliable models must identify the trustworthy sourc…

cs.LG2026

Not Every Divergence Should Be Suppressed: Counterfactual Recoverability in On-Policy Distillation

De Jiang, Zhengyang Zhang, Kehong Yuan +1

On-policy distillation (OPD) supervises student-visited trajectories, yet divergence-based rules cannot determine whether an erroneous prefix remains correctable. We formulate this…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.