3 papers
cs.CL2026
Mil-SCORE: Benchmarking Long-Context Geospatial Reasoning and Planning in Large Language Models
Aadi Palnitkar, Mingyang Mao, Nicholas Waytowich +2
As large language models (LLMs) are applied to increasingly longer and more complex tasks, there is a growing need for realistic long-context benchmarks that require selective read…
cs.HC2025
"New" Challenges for Future C2: Commanding Soldier-Machine Partnerships
Anna Madison, Kaleb McDowell, Vinicius G. Goecks +7
Future warfare will occur in more complex, fast-paced, ill-structured, and demanding conditions that will stress current Command and Control (C2) systems. Without modernization, th…
cs.AI2024
Atari-GPT: Benchmarking Multimodal Large Language Models as Low-Level Policies in Atari Games
Nicholas R. Waytowich, Devin White, MD Sunbeam +1
Recent advancements in large language models (LLMs) have expanded their capabilities beyond traditional text-based tasks to multimodal domains, integrating visual, auditory, and te…