3 papers
cs.HC2026
An LLM-Native Psychometric Instrument Reveals a Self-Report--Behavior Gap Across 25 Models
Juan Manuel Contreras
Large language models (LLMs) give stable answers to personality questionnaires, yet these self-reports fail to predict how the models behave. Is this gap an artifact of forcing hum…
cs.CV2025
Automated Evaluation of Gender Bias Across 13 Large Multimodal Models
Juan Manuel Contreras
Large multimodal models (LMMs) have revolutionized text-to-image generation, but they risk perpetuating the harmful social biases in their training data. Prior work has identified…
cs.AI2025
Policy-Grounded Safety Evaluation of 20 Large Language Models
Juan Manuel Contreras
As large language models (LLMs) become increasingly integrated into real-world applications, scalable and rigorous safety evaluation is essential. This paper introduces Aymara AI,…