Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Adversarial Reinforcement Learning for Large Language Model Agent Safety
Zizhao Wang, Dingcheng Li, Vaishakh Keshava +4
Large Language Model (LLM) agents can leverage tools such as Google Search to complete complex tasks. However, this tool usage introduces the risk of indirect prompt injections, wh…
cs.LG2025
Dyn-O: Building Structured World Models with Object-Centric Representations
Zizhao Wang, Kaixin Wang, Li Zhao +2
World models aim to capture the dynamics of the environment, enabling agents to predict and plan for future states. In most scenarios of interest, the dynamics are highly centered…