2 citations · 5 across the 13 of their papers we have counts for
16 papers · 1 filter
CrowdCue: Specialist-Cue Conditioning for Vision-Language Crowd Counting
Moshiur Farazi, Bekir Ciftler, Abdulhalim Dandoush +1
Generative vision-language models (VLMs) offer a counting paradigm in which one model produces both a count and a natural-language account of the scene, yet their raw counting accu…
When Do VLMs Help Arabic Manuscript OCR? A Cross-Dataset Study
Moshiur Farazi, Firoj Alam, Abderrahmane Maaradji +3
Vision-language models (VLMs) are increasingly being used for document understanding, yet their role in Arabic and Islamic manuscript recognition remains underexplored. To address…
The Gate Always Closes: On Injecting Auxiliary Signals into Frozen Vision-Language Models
Moshiur Farazi, Sameera Ramasinghe, Bekir Sait Ciftler +2
Auxiliary signal pathways in VLMs are routinely fitted with learnable gates so the optimiser can decide how much of the signal to admit. We find that the optimiser almost always de…
HyperVis: Continuous Latent Visual Relational Graphs on the Lorentz Hyperboloid for Compositional Reasoning
Moshiur Farazi, Sameera Ramasinghe, Mahbub Ahmed Turza +1
Vision-Language Models (VLMs) struggle with compositional reasoning that requires understanding inter-object relationships. A natural remedy is to inject explicit scene graph tripl…
Beyond the Pipeline: Analyzing Key Factors in End-to-End Deep Learning for Historical Writer Identification
Hanif Rasyidi, Moshiur Farazi
This paper investigates various factors that influence the performance of end-to-end deep learning approaches for historical writer identification (HWI), a task that remains challe…
Label Semantics for Robust Hyperspectral Image Classification
Rafin Hassan, Zarin Tasnim Roshni, Rafiqul Bari +4
Hyperspectral imaging (HSI) classification is a critical tool with widespread applications across diverse fields such as agriculture, environmental monitoring, medicine, and materi…