9 papers
When Do VLMs Help Arabic Manuscript OCR? A Cross-Dataset Study
Moshiur Farazi, Firoj Alam, Abderrahmane Maaradji +3
Vision-language models (VLMs) are increasingly being used for document understanding, yet their role in Arabic and Islamic manuscript recognition remains underexplored. To address…
The Gate Always Closes: On Injecting Auxiliary Signals into Frozen Vision-Language Models
Moshiur Farazi, Sameera Ramasinghe, Bekir Sait Ciftler +2
Auxiliary signal pathways in VLMs are routinely fitted with learnable gates so the optimiser can decide how much of the signal to admit. We find that the optimiser almost always de…
HyperVis: Continuous Latent Visual Relational Graphs on the Lorentz Hyperboloid for Compositional Reasoning
Moshiur Farazi, Sameera Ramasinghe, Mahbub Ahmed Turza +1
Vision-Language Models (VLMs) struggle with compositional reasoning that requires understanding inter-object relationships. A natural remedy is to inject explicit scene graph tripl…
Beyond the Pipeline: Analyzing Key Factors in End-to-End Deep Learning for Historical Writer Identification
Hanif Rasyidi, Moshiur Farazi
This paper investigates various factors that influence the performance of end-to-end deep learning approaches for historical writer identification (HWI), a task that remains challe…
Label Semantics for Robust Hyperspectral Image Classification
Rafin Hassan, Zarin Tasnim Roshni, Rafiqul Bari +4
Hyperspectral imaging (HSI) classification is a critical tool with widespread applications across diverse fields such as agriculture, environmental monitoring, medicine, and materi…
Multi-Modal Sentiment Analysis with Dynamic Attention Fusion
Sadia Abdulhalim, Muaz Albaghdadi, Moshiur Farazi
Traditional sentiment analysis has long been a unimodal task, relying solely on text. This approach overlooks non-verbal cues such as vocal tone and prosody that are essential for…