collaborators

13 papers

cs.CV2026

MorphUNet: Alpha-Controlled Biometric Transport for Diffusion-Based Face Morphing Attacks

Taimoor Rizwan, Sara Atito, Zhenhua Feng +2

Face morphing attacks create synthetic images verifiable against multiple identities, threatening border control and identity verification systems. We introduce MorphUNet, a diffus…

cs.CV2026

Diff-ID: Identity Consistent Facial Image Generation and Morphing via Diffusion Models

Taimoor Rizwan, Sara Atito, Muhammad Awais +2

Generative diffusion models have revolutionized facial image synthesis, yet robust identity preservation in high resolution outputs remains a critical challenge. This issue is espe…

cs.CV2026

CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks

Wish Suharitdamrong, Tony Alex, Muhammad Awais +1

Foundation models have revolutionized AI, but adapting them efficiently for multimodal tasks, particularly in dual-stream architectures composed of unimodal encoders, such as DINO…

cs.AI2026

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs

Wish Suharitdamrong, Muhammad Awais, Xiatian Zhu +1

Multimodal Large Language Models (MLLMs) can listen and see, but how do audio and visual signals actually travel through the network to shape an answer? Despite their growing role…

cs.LG2026

Information theoretic underpinning of self-supervised learning by clustering

Josef Kittler, Sara Atito, Muhammad Awais

Self-supervised learning (SSL) is recognized as an essential tool for building foundation models for Artificial Intelligence applications. The advances in SSL have been made thanks…

cs.CV2026

See Fair, Speak Truth: Equitable Attention Improves Grounding and Reduces Hallucination in Vision-Language Alignment

Mohammad Anas Azeez, Ankan Deria, Zohaib Hasan Siddiqui +5

Multimodal large language models (MLLMs) frequently hallucinate objects that are absent from the visual input, often because attention during decoding is disproportionately drawn t…