3 papers
cs.AI2026
LiteSearch-VL: Small Multimodal Search Agents via Trajectory Distillation and Synthetic Step-DPO
Saeed Khaki, Nima Safaei, Kamal Ginotra
Multimodal search agents answer visual questions by interleaving image understanding, web retrieval, tool use, and evidence synthesis. Strong systems exist, but in two expensive re…
cs.CV2026
Understanding Pruning Regimes in Vision-Language Models Through Domain-Aware Layer Selection
Saeed Khaki, Nima Safaei, Kamal Ginotra
Transformer-based vision-language models (VLMs) contain substantial depth redundancy, yet the effect of removing specific decoder layers remains poorly understood, especially for d…
cs.AI2026
VisTIRA: Closing the Image-Text Modality Gap in Visual Math Reasoning via Structured Tool Integration
Saeed Khaki, Ashudeep Singh, Nima Safaei +1
Vision-language models (VLMs) lag behind text-only language models on mathematical reasoning when the same problems are presented as images rather than text. We empirically charact…