2 papers
cs.CY2025
From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law
John Mavi, Diana Teodora Găitan, Sergio Coronado
Large Language Models (LLMs) are widely used across sectors, yet their alignment with International Humanitarian Law (IHL) is not well understood. This study evaluates eight leadin…
cs.CY2024
Assessing the Performance of Human-Capable LLMs -- Are LLMs Coming for Your Job?
John Mavi, Nathan Summers, Sergio Coronado
The current paper presents the development and validation of SelfScore, a novel benchmark designed to assess the performance of automated Large Language Model (LLM) agents on help…