3 papers
cs.CL2025
Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization
Masahiro Kaneko, Zeerak Talat, Timothy Baldwin
Iterative jailbreak methods that repeatedly rewrite and input prompts into large language models (LLMs) to induce harmful outputs -- using the model's previous responses to guide e…
cs.CR2025
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
Masahiro Kaneko, Timothy Baldwin
Adversarial attacks by malicious users that threaten the safety of large language models (LLMs) can be viewed as attempts to infer a target property that is unknown when an ins…
cs.CL2025
Balanced Multi-Factor In-Context Learning for Multilingual Large Language Models
Masahiro Kaneko, Alham Fikri Aji, Timothy Baldwin
Multilingual large language models (MLLMs) are able to leverage in-context learning (ICL) to achieve high performance by leveraging cross-lingual knowledge transfer without paramet…