Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Towards a Belief-Based World Model for LLM Agents
Shubham Kumar, Harshit Kumar, Narendra Ahuja +1
Large language models (LLMs) are being used as policies for autonomous decision-making and planning in many domains. Despite their strong reasoning capabilities, LLMs struggle with…
cs.AI2026
Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models
Shubham Kumar, Narendra Ahuja
Safety trained large language models (LLMs) can often be induced to answer harmful requests through jailbreak prompts. Because we lack a robust understanding of why LLMs are suscep…