209 citations · 322 across the 38 of their papers we have counts for
49 papers · 1 filter
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP
Sicheng Zhang, Muzammal Naseer, Binzhu Xie +5
CLIP and its variants are widely adopted visual backbones in multimodal systems, but their pretraining remains dominated by descriptive image-text alignment. As downstream applicat…
D3Seg: Dependency-Aware Diffusion for Brain Tumor Segmentation with Missing Modalities
Danish Ali, Ajmal Mian, Naveed Akhtar +1
Accurate brain tumor segmentation using multi-parametric MRI is critical for effective treatment planning. However, in clinical settings, complete acquisition of all MRI sequences…
DRBD-Mamba for Robust and Efficient Brain Tumor Segmentation with Analytical Insights
Danish Ali, Ajmal Mian, Naveed Akhtar +1
Accurate brain tumor segmentation is significant for clinical diagnosis and treatment but remains challenging due to tumor heterogeneity. Mamba-based State Space Models have demons…
CymbaDiff: Structured Spatial Diffusion for Sketch-based 3D Semantic Urban Scene Generation
Li Liang, Bo Miao, Xinyu Wang +3
Outdoor 3D semantic scene generation produces realistic and semantically rich environments for applications such as urban simulation and autonomous driving. However, advances in th…
On the Reliability of Vision-Language Models Under Adversarial Frequency-Domain Perturbations
Jordan Vice, Naveed Akhtar, Yansong Gao +2
Vision-Language Models (VLMs) are increasingly used as perceptual modules for visual content reasoning, including through captioning and DeepFake detection. In this work, we expose…
Multistream Network for LiDAR and Camera-based 3D Object Detection in Outdoor Scenes
Muhammad Ibrahim, Naveed Akhtar, Haitian Wang +2
Fusion of LiDAR and RGB data has the potential to enhance outdoor 3D object detection accuracy. To address real-world challenges in outdoor 3D object detection, fusion of LiDAR and…