3 papers
cs.CV2026
LLaVA-LE: Large Language-and-Vision Assistant for Lunar Exploration
Gokce Inal, Pouyan Navard, Alper Yilmaz
Recent advances in multimodal vision-language models (VLMs) have enabled joint reasoning over visual and textual information, yet their application to planetary science remains lar…
q-bio.QM2026
ERDES: A Benchmark Video Dataset for Retinal Detachment and Macular Status Classification in Ocular Ultrasound
Yasemin Ozkut, Pouyan Navard, Srikar Adhikari +4
Retinal detachment (RD) is a vision-threatening condition that requires prompt intervention to preserve sight. A critical factor in treatment urgency and visual prognosis is macula…
cs.CV2025
KnobGen: Controlling the Sophistication of Artwork in Sketch-Based Diffusion Models
Pouyan Navard, Amin Karimi Monsefi, Mengxi Zhou +3
Recent advances in diffusion models have significantly improved text-to-image (T2I) generation, but they often struggle to balance fine-grained precision with high-level control. M…