5 papers · 1 filter
AdaptPrompt: Parameter-Efficient Adaptation of VLMs for Generalizable Deepfake Detection
Yichen Jiang, Mohammed Talha Alam, Sohail Ahmed Khan +2
Detectors of AI-generated images tend to inherit the biases of the data they are trained on: models fitted to GAN imagery learn to treat GAN-specific artifacts as the very definiti…
PosA-VLA: Enhancing Action Generation via Pose-Conditioned Anchor Attention
Ziwen Li, Xin Wang, Hanlue Zhang +8
The Vision-Language-Action (VLA) models have demonstrated remarkable performance on embodied tasks and shown promising potential for real-world applications. However, current VLAs…
FaceAnonyMixer: Cancelable Faces via Identity Consistent Latent Space Mixing
Mohammed Talha Alam, Fahad Shamshad, Fakhri Karray +1
Advancements in face recognition (FR) technologies have amplified privacy concerns, necessitating methods that protect identity while maintaining recognition utility. Existing face…
UrbanGS: Semantic-Guided Gaussian Splatting for Urban Scene Reconstruction
Ziwen Li, Jiaxin Huang, Runnan Chen +5
Reconstructing urban scenes is challenging due to their complex geometries and the presence of potentially dynamic objects. 3D Gaussian Splatting (3DGS)-based methods have shown st…
CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging
Raza Imam, Mohammed Talha Alam, Umaima Rahman +2
Existing vision-text contrastive learning models enhance representation transferability and support zero-shot prediction by matching paired image and caption embeddings while pushi…