2 papers
cs.CL2026
Pragmatic Attack Surface: Vulnerabilities of Implicit Context in Large Language Models
Bocheng Chen, Han Zi, Roucheng Ou +5
In the era of large language models (LLMs), attackers often manipulate natural language to elicit unsafe or harmful outputs, creating a new natural language attack surface unique t…
cs.AI2025
Calibrating Transformer Attention via Task-Space Sensitivity Feedback
Yawei Liu
Transformer-based pre-trained language models (PLMs) excel in text classification but suffer from attention dilution and attention sink effects, forcing models to over-focus on tas…