3 papers
cs.CV2026
A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval
Mahyar Ghazanfari, Amin Tabrizian, Arsyi Aziz +2
Contrastive vision-language models such as CLIP and BLIP are typically trained on short image captions, limiting their ability to retrieve images from detailed textual descriptions…
cs.RO2026
End-to-End LLM Flight Planning with RAG-based Memory and Multi-modal Coach Agent
Amin Tabrizian, Arsyi Aziz, Aarifah Ullah +3
Bridging the gap between human pilot intent and autonomous flight operation is critical for real-world electric vertical takeoff and landing (eVTOL) aircraft deployment. Flight pla…
cs.RO2026
Transformer-based Multi-agent Reinforcement Learning for Separation Assurance in Structured and Unstructured Airspaces
Arsyi Aziz, Peng Wei
Conventional optimization-based metering depends on strict adherence to precomputed schedules, which limits the flexibility required for the stochastic operations of Advanced Air M…