Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Sockpuppetting: Jailbreaking LLMs by Combining Prefilling with Optimization
Asen Dotsinski, Panagiotis Eustratiadis
Prefill attacks are an effective and low-cost jailbreaking method, as they directly insert an acceptance sequence (e.g., "Sure, here is...") at the start of an LLM's output and lea…
cs.CL2025
On the Generalizability of "Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals"
Asen Dotsinski, Udit Thakur, Marko Ivanov +2
We present a reproduction study of "Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals" (Ortu et al., 2024), which investigates competition of…