4 papers
Uncertainty-Aware Decision Making in Multimodal Large Language Models
Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed
Multimodal large language models (MLLMs) increasingly answer questions whose correctness depends on visual, textual, temporal, acoustic, document, chart, or embodied evidence. Thei…
SegRAG: Training-Free Retrieval-Augmented Semantic Segmentation
Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed
Open-vocabulary segmentation models such as SAM3 perform well across broad categories via text prompting, yet degrade when target classes are visually underrepresented in pretraini…
CropVLM: A Domain-Adapted Vision-Language Model for Open-Set Crop Analysis
Abderrahmene Boudiaf, Sajd Javed
High-throughput plant phenotyping, the quantitative measurement of observable plant traits, is critical for modern breeding but remains constrained by a "phenotyping bottleneck," w…
AgriChat: A Multimodal Large Language Model for Agriculture Image Understanding
Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed
The deployment of Multimodal Large Language Models (MLLMs) in agriculture is currently stalled by a critical trade-off: the existing literature lacks the large-scale agricultural d…