2 papers
cs.AI2026
MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs
Albert Wu, Nicholas Roberts, Tzu-Heng Huang +5
LLM coding agents now generate complex programs at a scale that makes thorough human review increasingly difficult, raising the risk of safety and security failures. Common approac…
cs.LG2026
Test-Time Scaling Makes Overtraining Compute-Optimal
Nicholas Roberts, Sungjun Cho, Zhiqi Gao +7
Modern LLMs scale at test-time, e.g. via repeated sampling, where inference cost grows with model size and the number of samples. This creates a trade-off that pretraining scaling…