collaborators

5 papers

cs.CL2026

Shieldstral

Antonia Calvi, Avinash Sooriyarachchi, Giada Pistilli +274

We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7 its size on text safety benchmarks and set…

cs.SE2026

TerraFormer: Automated Infrastructure-as-Code with LLMs Fine-Tuned via Policy-Guided Verifier Feedback

Prithwish Jana, Sam Davidson, Bhavana Bhasker +3

Automating Infrastructure-as-Code (IaC) is challenging, and large language models (LLMs) often produce incorrect configurations from natural language (NL). We present TerraFormer,…

cs.DC2025

Multi-IaC-Eval: Benchmarking Cloud Infrastructure as Code Across Multiple Formats

Sam Davidson, Li Sun, Bhavana Bhasker +2

Infrastructure as Code (IaC) is fundamental to modern cloud computing, enabling teams to define and manage infrastructure through machine-readable configuration files. However, dif…

cs.SE2025

SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents

Muhammad Shihab Rashid, Christian Bock, Yuan Zhuang +10

Coding agents powered by large language models have shown impressive capabilities in software engineering tasks, but evaluating their performance across diverse programming languag…

cs.CL2025

Large Language Model Critics for Execution-Free Evaluation of Code Changes

Aashish Yadavally, Hoan Nguyen, Laurent Callot +1

Large language models (LLMs) offer a promising way forward for automating software engineering tasks, such as bug fixes, feature additions, etc., via multi-step LLM-based agentic w…