1 citations · 1 across the 4 of their papers we have counts for
4 papers
PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs
Dayang Liang, Liyuan He, Xuan Feng +3
Group-relative policy optimization has emerged as a key paradigm for training agentic large language models (LLMs) on multi-turn interactive tasks. However, most existing variants…
Where Detectors Fail: Closing the Tail-Domain Gap with Expert-Guided Mutual Distillation
Xuan Feng, Guihong Liu, Tianlong Gu +5
Multimodal fake news detectors often generalize poorly across domains because they learn to trust unreliable evidence: domain-specific shortcuts amplified by imbalanced data and se…
C2PO: Diagnosing and Disentangling Bias Shortcuts in LLMs
Xuan Feng, Bo An, Tianlong Gu +4
Bias in Large Language Models (LLMs) poses significant risks to trustworthiness, manifesting primarily as stereotypical biases (e.g., gender or racial stereotypes) and structural b…
Learning from Mistakes: Self-correct Adversarial Training for Chinese Unnatural Text Correction
Xuan Feng, Tianlong Gu, Xiaoli Liu +1
Unnatural text correction aims to automatically detect and correct spelling errors or adversarial perturbation errors in sentences. Existing methods typically rely on fine-tuning o…