collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2025

BlackboxNLP-2025 MIB Shared Task: Exploring Ensemble Strategies for Circuit Localization Methods

Philipp Mondorf, Mingyang Wang, Sebastian Gerstner +6

The Circuit Localization track of the Mechanistic Interpretability Benchmark (MIB) evaluates methods for localizing circuits within large language models (LLMs), i.e., subnetworks…

cs.CL2025

Refusal Direction is Universal Across Safety-Aligned Languages

Xinpeng Wang, Mingyang Wang, Yihong Liu +2

Refusal mechanisms in large language models (LLMs) are essential for ensuring safety. Recent research has revealed that refusal behavior can be mediated by a single direction in ac…

cs.CL2025

Tracing Multilingual Factual Knowledge Acquisition in Pretraining

Yihong Liu, Mingyang Wang, Amir Hossein Kargaran +5

Large Language Models (LLMs) are capable of recalling multilingual factual knowledge present in their pretraining data. However, most studies evaluate only the final model, leaving…

cs.CL2025

Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior

Shengyun Si, Xinpeng Wang, Guangyao Zhai +2

Recent advancements in large language models (LLMs) have demonstrated that fine-tuning and human alignment can render LLMs harmless. In practice, such "harmlessness" behavior is ma…

cs.CL2025

M-ABSA: A Multilingual Dataset for Aspect-Based Sentiment Analysis

Chengyan Wu, Bolei Ma, Yihong Liu +7

Aspect-based sentiment analysis (ABSA) is a crucial task in information extraction and sentiment analysis, aiming to identify aspects with associated sentiment elements in text. Ho…

cs.CL2024

Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation

Xinpeng Wang, Chengzhi Hu, Paul Röttger +1

Training a language model to be both helpful and harmless requires careful calibration of refusal behaviours: Models should refuse to follow malicious instructions or give harmful…