Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
A Theoretical Game of Attacks via Compositional Skills
Xinbo Wu, Huan Zhang, Abhishek Umrawal +1
As large language models grow increasingly capable, concerns about their safe deployment have intensified. While numerous alignment strategies aim to restrict harmful behavior, the…
cs.CL2025
Concealment of Intent: A Game-Theoretic Analysis
Xinbo Wu, Abhishek Umrawal, Lav R. Varshney
As large language models (LLMs) grow more capable, concerns about their safe deployment have also grown. Although alignment mechanisms have been introduced to deter misuse, they re…
cs.CL2025
JAM: Controllable and Responsible Text Generation via Causal Reasoning and Latent Vector Manipulation
Yingbing Huang, Deming Chen, Abhishek K. Umrawal
While large language models (LLMs) have made significant strides in generating coherent and contextually relevant text, they often function as opaque black boxes, trained on vast u…