3 papers
cs.CL2026
Evaluating Autoformalization Robustness via Semantically Similar Paraphrasing
Hayden Moore, Asfahan Shah
Large Language Models (LLMs) have recently emerged as powerful tools for autoformalization. Despite their impressive performance, these models can still struggle to produce grounde…
cs.CL2026
A Pilot Benchmark for NL-to-FOL Translation in Planetary Exploration
Hayden Moore, Suman Saha, Mahfuza Farooque
Future planetary exploration envisions autonomous robotic agents operating under severe communication constraints, without global positioning, and with minimal human intervention.…
cs.CR2026
Trojans in Artificial Intelligence (TrojAI) Final Report
Kristopher W. Reese, Taylor Kulp-McDowall, Michael Majurski +68
The Intelligence Advanced Research Projects Activity (IARPA) launched the TrojAI program to confront an emerging vulnerability in modern artificial intelligence: the threat of AI T…