2 papers
cs.SE2026
ATime-Consistent Benchmark for Repository-Level Software Engineering Evaluation
Xianpeng, Sun, Haonan Sun +5
Evaluation of repository-aware software engineering systems is often confounded by synthetic task design, prompt leakage, and temporal contamination between repository knowledge an…
cs.SE2026
The Llama 4 Herd: Architecture, Training, Evaluation, and Deployment Notes
Redacted by arXiv
This document consolidates publicly reported technical details about Metas Llama 4 model family. It summarizes (i) released variants (Scout and Maverick) and the broader herd conte…