4 papers
Understanding the Self-Reflection Mechanisms of LLMs through Biased Attitude Associations
Jingshen Zhang, Bo Wang, Boci Yang +4
While the emergent self-reflection capabilities of Large Language Models (LLMs) offer a promising paradigm for autonomous bias mitigation, their internal mechanics remain unclear,…
Modeling Implicit Conflict Monitoring Mechanisms against Stereotypes in LLMs
Jingshen Zhang, Bo Wang, Yanlin Fu +4
In this paper, we study an emergent self-debiasing mechanisms against stereotypical content in Large Language Models (LLMs). Unlike traditional safety mechanisms that are primarily…
DualReward: A Dynamic Reinforcement Learning Framework for Cloze Tests Distractor Generation
Tianyou Huang, Xinglu Chen, Jingshen Zhang +2
This paper introduces DualReward, a novel reinforcement learning framework for automatic distractor generation in cloze tests. Unlike conventional approaches that rely primarily on…
Label Confidence Weighted Learning for Target-level Sentence Simplification
Xinying Qiu, Jingshen Zhang
Multi-level sentence simplification generates simplified sentences with varying language proficiency levels. We propose Label Confidence Weighted Learning (LCWL), a novel approach…