4 papers · 1 filter
Shieldstral
Antonia Calvi, Avinash Sooriyarachchi, Giada Pistilli +274
We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7 its size on text safety benchmarks and set…
Ministral 3
Alexander H. Liu, Kartik Khandelwal, Sandeep Subramanian +116
We introduce the Ministral 3 series, a family of parameter-efficient dense language models designed for compute and memory constrained applications, available in three model sizes:…
Large Language Model Critics for Execution-Free Evaluation of Code Changes
Aashish Yadavally, Hoan Nguyen, Laurent Callot +1
Large language models (LLMs) offer a promising way forward for automating software engineering tasks, such as bug fixes, feature additions, etc., via multi-step LLM-based agentic w…
Automated Evaluation of Retrieval-Augmented Language Models with Task-Specific Exam Generation
Gauthier Guinet, Behrooz Omidvar-Tehrani, Anoop Deoras +1
We propose a new method to measure the task-specific accuracy of Retrieval-Augmented Large Language Models (RAG). Evaluation is performed by scoring the RAG on an automatically-gen…