4 papers · 1 filter
A Theoretical Game of Attacks via Compositional Skills
Xinbo Wu, Huan Zhang, Abhishek Umrawal +1
As large language models grow increasingly capable, concerns about their safe deployment have intensified. While numerous alignment strategies aim to restrict harmful behavior, the…
Concealment of Intent: A Game-Theoretic Analysis
Xinbo Wu, Abhishek Umrawal, Lav R. Varshney
As large language models (LLMs) grow more capable, concerns about their safe deployment have also grown. Although alignment mechanisms have been introduced to deter misuse, they re…
SwitchCIT: Switching for Continual Instruction Tuning
Xinbo Wu, Max Hartman, Vidhata Arjun Jayaraman +1
Large language models (LLMs) and multimodal models (MMs) have exhibited impressive capabilities in various domains, particularly in general language understanding and visual reason…
Transformer-based Causal Language Models Perform Clustering
Xinbo Wu, Lav R. Varshney
Even though large language models (LLMs) have demonstrated remarkable capability in solving various natural language tasks, the capability of an LLM to follow human instructions is…