5 papers
A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval
Mahyar Ghazanfari, Amin Tabrizian, Arsyi Aziz +2
Contrastive vision-language models such as CLIP and BLIP are typically trained on short image captions, limiting their ability to retrieve images from detailed textual descriptions…
End-to-End LLM Flight Planning with RAG-based Memory and Multi-modal Coach Agent
Amin Tabrizian, Arsyi Aziz, Aarifah Ullah +3
Bridging the gap between human pilot intent and autonomous flight operation is critical for real-world electric vertical takeoff and landing (eVTOL) aircraft deployment. Flight pla…
EO-Agents: A Three-Agent LLM Pipeline for Earth Observation Hypothesis Generation
Mahyar Ghazanfari, Amin Tabrizian, Armin Mehrabian +1
Large language models have recently been explored for scientific hypothesis generation, but most prior work relies on unstructured literature and free-form textual claims. We prese…
Delay-Aware Reinforcement Learning for Highway On-Ramp Merging under Stochastic Communication Latency
Amin Tabrizian, Zhitong Huang, Arsyi Aziz +1
Delayed and partially observable state information poses significant challenges for reinforcement learning (RL)-based control in real-world autonomous driving. In highway on-ramp m…
Towards Automated Air Traffic Safety Assessment Around Non-Towered Airports Using Large Language Models
Torsten Darrell, Mahyar Ghazanfari, Jordan Kam +3
We investigate frameworks for post-flight safety analysis at non-towered airports using large language models (LLMs). Non-towered airports rely on the Common Traffic Advisory Frequ…