collaborators

7 papers

cs.MA2026

An Actionable Diagnosis of Multilingual, Multi-Agent Planning Failures

Vikas Pahuja, Jonathan Brokman, Omer Hofman +6

Multilingual multi-agent systems exhibit substantial degradation beyond English, yet prior work rarely identifies how task-critical information is lost when user requests are conve…

cs.CL2026

Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types

Hadas Orgad, Boyi Wei, Kaden Zheng +4

Large language models (LLMs) undergo alignment training to avoid harmful behaviors, yet the resulting safeguards remain brittle: jailbreaks routinely bypass them, and fine-tuning o…

cs.AI2026

Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators

Anissa Alloula, Federico Licini, Ava Batchkala +1

LLMs-as-judges are the only way to evaluate safety at scale. Despite their importance, LLM-judges themselves are rarely evaluated beyond human agreement in simple, static benchmark…

cs.LG2026

Learning is Forgetting: LLM Training As Lossy Compression

Henry C. Conklin, Tom Hosking, Tan Yi-Chern +5

Despite the increasing prevalence of large language models (LLMs), we still have a limited understanding of how their representational spaces are structured. This limits our abilit…

cs.HC2026

Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations

Preethi Seshadri, Samuel Cahyawijaya, Ayomide Odumakinde +2

Agentic benchmarks increasingly rely on LLM-simulated users to scalably evaluate agent performance, yet the robustness, validity, and fairness of this approach remain unexamined. T…

cs.CL2025

A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users

Nishant Balepur, Matthew Shu, Yoo Yeon Sung +5

To assist users in complex tasks, LLMs generate plans: step-by-step instructions towards a goal. While alignment methods aim to ensure LLM plans are helpful, they train (RLHF) or e…