◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Karolina Korgul

5 papers hereh-index 373 citations9 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author4

Across the 5 of 5 papers where every author was matched, so the position is known.

fields
  • cs.CL3
  • cs.CY1
  • cs.HC1

identity via Semantic Scholar / OpenAlex

activity
20232026
most citedIt's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents

1 citations · 1 across the 2 of their papers we have counts for

collaborators
Showing cs.CLShow all

3 papers · 1 filter

cs.CL2025

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou +39

Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstrac…

cs.CL2025

LINGOLY-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic Obfuscation

Jude Khouja, Lingyi Yang, Karolina Korgul +6

Frontier language models demonstrate increasing ability at solving reasoning problems, but their performance is often inflated by circumventing reasoning and instead relying on the…

cs.CL2023

Do Large Language Models have Shared Weaknesses in Medical Question Answering?

Andrew M. Bean, Karolina Korgul, Felix Krones +2

Large language models (LLMs) have made rapid improvement on medical benchmarks, but their unreliability remains a persistent challenge for safe real-world uses. To design for the u…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.