3 papers
cs.CV2026
Towards a Large Language-Vision Question Answering Model for MSTAR Automatic Target Recognition
David F. Ramirez, Tim L. Overman, Kristen Jaskie +2
Large language-vision models (LLVM), such as OpenAI's ChatGPT and GPT-4, have gained prominence as powerful tools for analyzing text and imagery. The merging of these data domains…
eess.IV2026
Geospatial-Temporal Sensemaking of Remote Sensing Activity Detections with Multimodal Large Language Model
David F. Ramirez, Tim Overman, Kristen Jaskie +1
We introduce SMART-HC-VQA, a Sentinel-2-based visual question answering dataset derived from the IARPA SMART Heavy Construction dataset, designed for spatiotemporal analysis of hum…
cs.CV2026
SAR-RAG: ATR Visual Question Answering by Semantic Search, Retrieval, and MLLM Generation
David F. Ramirez, Tim Overman, Kristen Jaskie +2
We present a visual-context image-retrieval-augmented generation (ImageRAG)- assisted AI agent for automatic target recognition (ATR) of synthetic aperture radar (SAR) imagery. SAR…