6 papers
Metamorphic Testing of Vision-Language Action-Enabled Robots
Pablo Valle, Sergio Segura, Shaukat Ali +1
Vision-Language-Action (VLA) models are multimodal robotic task controllers that, given an instruction and visual inputs, produce a sequence of low-level control actions (or motor…
Meta-Fair: AI-Assisted Fairness Testing of Large Language Models
Miguel Romero-Arjona, José A. Parejo, Juan C. Alonso +3
Fairness--the absence of unjustified bias--is a core principle in the development of Artificial Intelligence (AI) systems, yet it remains difficult to assess and enforce. Current a…
Red Teaming Contemporary AI Models: Insights from Spanish and Basque Perspectives
Miguel Romero-Arjona, Pablo Valle, Juan C. Alonso +7
The battle for AI leadership is on, with OpenAI in the United States and DeepSeek in China as key contenders. In response to these global trends, the Spanish government has propose…
o3-mini vs DeepSeek-R1: Which One is Safer?
Aitor Arrieta, Miriam Ugarte, Pablo Valle +2
The irruption of DeepSeek-R1 constitutes a turning point for the AI industry in general and the LLMs in particular. Its capabilities have demonstrated outstanding performance in se…
Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation
Aitor Arrieta, Miriam Ugarte, Pablo Valle +2
Large Language Models (LLMs) have become an integral part of our daily lives. However, they impose certain risks, including those that can harm individuals' privacy, perpetuate bia…
ASTRAL: Automated Safety Testing of Large Language Models
Miriam Ugarte, Pablo Valle, José Antonio Parejo +2
Large Language Models (LLMs) have recently gained attention due to their ability to understand and generate sophisticated human-like content. However, ensuring their safety is para…