activity
20242026
collaborators

8 papers

cs.DB2026

MAPS: A Multilingual Benchmark for Agent Performance and Security

Omer Hofman, Jonathan Brokman, Oren Rachmil +7

Agentic AI systems, which build on Large Language Models (LLMs) and interact with tools and memory, have rapidly advanced in capability and scope. Yet, since LLMs have been shown t…

cs.CL2025

Small Changes, Large Consequences: Analyzing the Allocational Fairness of LLMs in Hiring Contexts

Preethi Seshadri, Hongyu Chen, Sameer Singh +1

Large language models (LLMs) are increasingly being deployed in high-stakes applications like hiring, yet their potential for unfair decision-making remains understudied in generat…

cs.CL2025

Safer or Luckier? LLMs as Safety Evaluators Are Not Robust to Artifacts

Hongyu Chen, Seraphina Goldfarb-Tarrant

Large Language Models (LLMs) are increasingly employed as automated evaluators to assess the safety of generated content, yet their reliability in this role remains uncertain. This…

cs.AI2025

The Multilingual Divide and Its Impact on Global AI Safety

Aidan Peppin, Julia Kreutzer, Alice Schoenauer Sebag +13

Despite advances in large language model capabilities in recent years, a large gap remains in their capabilities and safety performance for many languages beyond a relatively small…

cs.CL2025

Command A: An Enterprise-Ready Large Language Model

Team Cohere, :, Aakanksha +227

In this report we describe the development of Command A, a powerful large language model purpose-built to excel at real-world enterprise use cases. Command A is an agent-optimised…

cs.CL2024

Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning

Aakanksha, Arash Ahmadian, Seraphina Goldfarb-Tarrant +3

Large Language Models (LLMs) have been adopted and deployed worldwide for a broad variety of applications. However, ensuring their safe use remains a significant challenge. Prefere…