activity
20242026
collaborators

5 papers

cs.AI2026

Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models

Jasmine Brazilek, Joel Christoph, Maheep Chaudhary +4

Previous research has evaluated animal welfare using question-and-answer benchmarks. This study investigates whether these evaluations also hold in agentic settings. The agents may…

cs.CY2026

Prioritization of Risks from Artificial Intelligence: A Delphi Study of 272 International Experts

Alexander K. Saeri, Jess Graham, Michael Noetel +185

Artificial intelligence poses many risks, ranging from familiar present-day harms to unprecedented and potentially catastrophic ones. Effective risk management requires prioritizat…

cs.CY2025

What do Large Language Models Say About Animals? Investigating Risks of Animal Harm in Generated Text

Arturs Kanepajs, Aditi Basu, Sankalpa Ghose +7

As machine learning systems become increasingly embedded in society, their impact on human and nonhuman life continues to escalate. Technical evaluations have addressed a variety o…

cs.CL2025

LAG-MMLU: Benchmarking Frontier LLM Understanding in Latvian and Giriama

Naome A. Etori, Kevin Lu, Randu Karisa +1

As large language models (LLMs) rapidly advance, evaluating their performance is critical. LLMs are trained on multilingual data, but their reasoning abilities are mainly evaluated…

cs.CL2024

Towards Safe Multilingual Frontier AI

Artūrs Kanepajs, Vladimir Ivanov, Richard Moulange

Linguistically inclusive LLMs -- which maintain good performance regardless of the language with which they are prompted -- are necessary for the diffusion of AI benefits around th…