1 citations · 1 across the 3 of their papers we have counts for
8 papers
It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents
Karolina Korgul, Yushi Yang, Arkadiusz Drohomirecki +7
Web-based agents powered by large language models are increasingly used for tasks such as email management or professional networking. Their reliance on dynamic web content, howeve…
RepSelect: Robust LLM Unlearning via Representation Selectivity
Filip Sondej, Yushi Yang, Adam Mahdi
When LLM weights are open or fine-tuning is available through an API, suppressing hazardous knowledge and tendencies is not enough: removal has to be deep enough that an adversary…
Agentic Reinforcement Learning for Search Misaligns Instruction-Tuning
Yushi Yang, Shreyansh Padarha, Sarah Ball +2
Agentic reinforcement learning (RL) trains large language models to use tools, but its impact on alignment is poorly understood. We study how agentic RL for search affects the alig…
LINGOLY-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic Obfuscation
Jude Khouja, Lingyi Yang, Karolina Korgul +6
Frontier language models demonstrate increasing ability at solving reasoning problems, but their performance is often inflated by circumventing reasoning and instead relying on the…
Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections
Åukasz Borchmann, Jordy Van Landeghem, MichaÅ Turski +12
Multimodal agents offer a promising path to automating complex document-intensive workflows. Yet, a critical question remains: do these agents demonstrate genuine strategic reasoni…
A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior
Harry Mayne, Justin Singh Kang, Dewi Gould +3
LLM self-explanations are often presented as a promising tool for AI oversight, yet their faithfulness to the model's true reasoning process is poorly understood. Existing faithful…