2 papers
cs.MA2026
An Actionable Diagnosis of Multilingual, Multi-Agent Planning Failures
Vikas Pahuja, Jonathan Brokman, Omer Hofman +6
Multilingual multi-agent systems exhibit substantial degradation beyond English, yet prior work rarely identifies how task-critical information is lost when user requests are conve…
cs.CR2026
When Scanners Lie: Evaluator Instability in LLM Red-Teaming
Lidor Erez, Omer Hofman, Tamir Nizri +1
Automated LLM vulnerability scanners are increasingly used to assess security risks by measuring different attack type success rates (ASR). Yet the validity of these measurements h…