activity
20242026
most citedIt's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents

1 citations · 1 across the 2 of their papers we have counts for

collaborators

5 papers

cs.HC20261 cited

It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents

Karolina Korgul, Yushi Yang, Arkadiusz Drohomirecki +7

Web-based agents powered by large language models are increasingly used for tasks such as email management or professional networking. Their reliance on dynamic web content, howeve…

cs.CL2026

LINGOLY-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic Obfuscation

Jude Khouja, Lingyi Yang, Karolina Korgul +6

Frontier language models demonstrate increasing ability at solving reasoning problems, but their performance is often inflated by circumventing reasoning and instead relying on the…

cs.CY2026

Agent Benchmarks Fail Public Sector Requirements

Jonathan Rystrøm, Chris Schmitz, Karolina Korgul +2

Deploying Large Language Model-based agents (LLM agents) in the public sector requires assuring that they meet the stringent legal, procedural, and structural requirements of publi…

cs.CL2025

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou +39

Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstrac…

cs.CL2024

Do Large Language Models have Shared Weaknesses in Medical Question Answering?

Andrew M. Bean, Karolina Korgul, Felix Krones +2

Large language models (LLMs) have made rapid improvement on medical benchmarks, but their unreliability remains a persistent challenge for safe real-world uses. To design for the u…