Publications (55)
SSR-Zero: Simple Self-Rewarding Reinforcement Learning for Machine Translation
Wenjie Yang, Mao Zheng, Mingyang Song +2
Large language models (LLMs) have recently demonstrated remarkable capabilities in machine translation (MT). However, most advanced MT-specific LLMs heavily rely on external superv…
PRISM: Probability Reallocation with In-Span Masking for Knowledge-Sensitive Alignment
Chenning Xu, Mao Zheng, Mingyang Song
Supervised fine-tuning (SFT) with token-level hard labels can amplify overconfident imitation of factually unsupported targets, causing hallucinations that propagate in multi-sente…
Hyperbolic Relevance Matching for Neural Keyphrase Extraction
Mingyang Song, Yi Feng, Liping Jing
Keyphrase extraction is a fundamental task in natural language processing and information retrieval that aims to extract a set of phrases with important information from a source d…
Importance Estimation from Multiple Perspectives for Keyphrase Extraction
Mingyang Song, Liping Jing, Lin Xiao
Keyphrase extraction is a fundamental task in Natural Language Processing, which usually contains two main parts: candidate keyphrase extraction and keyphrase importance estimation…
PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails
Mingyang Song, Luxin Xu, Haoyu Sun +3
Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an intrinsic property of an image. Real deployments are different: t…
Rubric-based On-policy Distillation
Junfeng Fang, Zhepei Hong, Mao Zheng +7
On-policy distillation (OPD) is a powerful paradigm for model alignment, yet its reliance on teacher logits restricts its application to white-box scenarios. We contend that struct…