19 citations · 24 across the 18 of their papers we have counts for
19 papers · 1 filter
From Gaze to Meaning: A Training-Free AI Agent for Unified Grounding and Explanation
Shayan Nasiriboukani, Sara Atito, Mohammad Nezamipour +1
Understanding human attention is fundamental for scene interpretation, yet existing approaches often rely on heavily trained models that lack interpretability. Prior methods strugg…
MorphUNet: Alpha-Controlled Biometric Transport for Diffusion-Based Face Morphing Attacks
Taimoor Rizwan, Sara Atito, Zhenhua Feng +2
Face morphing attacks create synthetic images verifiable against multiple identities, threatening border control and identity verification systems. We introduce MorphUNet, a diffus…
Diff-ID: Identity Consistent Facial Image Generation and Morphing via Diffusion Models
Taimoor Rizwan, Sara Atito, Muhammad Awais +2
Generative diffusion models have revolutionized facial image synthesis, yet robust identity preservation in high resolution outputs remains a critical challenge. This issue is espe…
See Fair, Speak Truth: Equitable Attention Improves Grounding and Reduces Hallucination in Vision-Language Alignment
Mohammad Anas Azeez, Ankan Deria, Zohaib Hasan Siddiqui +5
Multimodal large language models (MLLMs) frequently hallucinate objects that are absent from the visual input, often because attention during decoding is disproportionately drawn t…
CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks
Wish Suharitdamrong, Tony Alex, Muhammad Awais +1
Foundation models have revolutionized AI, but adapting them efficiently for multimodal tasks, particularly in dual-stream architectures composed of unimodal encoders, such as DINO…
Domain Adaptation Without the Compute Burden for Efficient Whole Slide Image Analysis
Umar Marikkar, Muhammad Awais, Sara Atito
Computational methods on analyzing Whole Slide Images (WSIs) enable early diagnosis and treatments by supporting pathologists in detection and classification of tumors. However, th…