2 papers
cs.LG2026
Detection of adversarial intent in Human-AI teams using LLMs
Abed K. Musaffar, Ambuj Singh, Francesco Bullo
Large language models (LLMs) are increasingly deployed in human-AI teams as support agents for complex tasks such as information retrieval, programming, and decision-making assista…
cs.HC2025
Learning to Lie: Reinforcement Learning Attacks Damage Human-AI Teams and Teams of LLMs
Abed Kareem Musaffar, Anand Gokhale, Sirui Zeng +4
As artificial intelligence (AI) assistants become more widely adopted in safety-critical domains, it becomes important to develop safeguards against potential failures or adversari…