2 citations · 2 across the 6 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
HAMSA: Hijacking Aligned Compact Models via Stealthy Automation
Alexey Krylov, Iskander Vagizov, Dmitrii Korzh +6
Large Language Models (LLMs), especially their compact efficiency-oriented variants, remain susceptible to jailbreak attacks that can elicit harmful outputs despite extensive align…
cs.CL2025★ 2 cited
Geopolitical biases in LLMs: what are the "good" and the "bad" countries according to contemporary language models
Mikhail Salnikov, Dmitrii Korzh, Ivan Lazichny +7
This paper evaluates geopolitical biases in LLMs with respect to various countries though an analysis of their interpretation of historical events with conflicting national perspec…