collaborators

5 papers

cs.CL2026

Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator

Shiping Yang, Shining Liang, Weihao Liu +4

Identifying faithfulness hallucinations in LLM-generated outputs remains challenging due to the scarcity of high-quality annotated data. Recent work relies on advanced LLMs to synt…

cs.CL2026

Quantifying and Improving the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data

Shiping Yang, Jie Wu, Wenbiao Ding +7

Robustness has become a critical attribute for the deployment of RAG systems in real-world applications. Existing research focuses on robustness to explicit noise (e.g., document s…

cs.CL2026

PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratch

Shangjian Yin, Shining Liang, Wenbiao Ding +4

High-quality instruction data is critical for LLM alignment, yet existing open-source datasets often lack efficiency, requiring hundreds of thousands of examples to approach propri…

cs.CL2025

Selected Languages are All You Need for Cross-lingual Truthfulness Transfer

Weihao Liu, Ning Wu, Wenbiao Ding +3

Truthfulness stands out as an essential challenge for Large Language Models (LLMs). Although many works have developed various ways for truthfulness enhancement, they seldom focus…

cs.CL2025

MuDAF: Long-Context Multi-Document Attention Focusing through Contrastive Learning on Attention Heads

Weihao Liu, Ning Wu, Shiping Yang +4

Large Language Models (LLMs) frequently show distracted attention due to irrelevant information in the input, which severely impairs their long-context capabilities. Inspired by re…