5 papers
Shieldstral
Antonia Calvi, Avinash Sooriyarachchi, Giada Pistilli +274
We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7 its size on text safety benchmarks and set…
TerraFormer: Automated Infrastructure-as-Code with LLMs Fine-Tuned via Policy-Guided Verifier Feedback
Prithwish Jana, Sam Davidson, Bhavana Bhasker +3
Automating Infrastructure-as-Code (IaC) is challenging, and large language models (LLMs) often produce incorrect configurations from natural language (NL). We present TerraFormer,…
Multi-IaC-Eval: Benchmarking Cloud Infrastructure as Code Across Multiple Formats
Sam Davidson, Li Sun, Bhavana Bhasker +2
Infrastructure as Code (IaC) is fundamental to modern cloud computing, enabling teams to define and manage infrastructure through machine-readable configuration files. However, dif…
SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents
Muhammad Shihab Rashid, Christian Bock, Yuan Zhuang +10
Coding agents powered by large language models have shown impressive capabilities in software engineering tasks, but evaluating their performance across diverse programming languag…
Large Language Model Critics for Execution-Free Evaluation of Code Changes
Aashish Yadavally, Hoan Nguyen, Laurent Callot +1
Large language models (LLMs) offer a promising way forward for automating software engineering tasks, such as bug fixes, feature additions, etc., via multi-step LLM-based agentic w…