16 papers · 1 filter
From Gaze to Meaning: A Training-Free AI Agent for Unified Grounding and Explanation
Shayan Nasiriboukani, Sara Atito, Mohammad Nezamipour +1
Understanding human attention is fundamental for scene interpretation, yet existing approaches often rely on heavily trained models that lack interpretability. Prior methods strugg…
MorphUNet: Alpha-Controlled Biometric Transport for Diffusion-Based Face Morphing Attacks
Taimoor Rizwan, Sara Atito, Zhenhua Feng +2
Face morphing attacks create synthetic images verifiable against multiple identities, threatening border control and identity verification systems. We introduce MorphUNet, a diffus…
Diff-ID: Identity Consistent Facial Image Generation and Morphing via Diffusion Models
Taimoor Rizwan, Sara Atito, Muhammad Awais +2
Generative diffusion models have revolutionized facial image synthesis, yet robust identity preservation in high resolution outputs remains a critical challenge. This issue is espe…
The Hidden Evolution of Disguised Visual Context inside the VLM
Wish Suharitdamrong, Tony Alex, Xiatian Zhu +2
Visual tokens enter Large Language Models (LLMs) as raw, foreign signals. How they are transformed into meaningful representations and interact with the language space depends enti…
Channel-Aware Probing for Multi-Channel Imaging
Umar Marikkar, Syed Sameed Husain, Muhammad Awais +1
Training and evaluating vision encoders on Multi-Channel Imaging (MCI) data remains challenging as channel configurations vary across datasets, preventing fixed-channel training an…
MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity Control
Fatemeh Nazarieh, Zhenhua Feng, Diptesh Kanojia +2
Audio-driven talking face generation has gained significant attention for applications in digital media and virtual avatars. While recent methods improve audio-lip synchronization,…