9 papers
KineVLA: Towards Kinematics-Aware Vision-Language-Action Models with Bi-Level Action Decomposition
Gaoge Han, Zhengqing Gao, Ziwen Li +5
In this paper, we introduce a novel kinematics-rich vision-language-action (VLA) task, in which language commands densely encode diverse kinematic attributes (such as direction, tr…
AdaptPrompt: Parameter-Efficient Adaptation of VLMs for Generalizable Deepfake Detection
Yichen Jiang, Mohammed Talha Alam, Sohail Ahmed Khan +2
Recent advances in image generation have led to the widespread availability of highly realistic synthetic media, increasing the difficulty of reliable deepfake detection. A key cha…
PosA-VLA: Enhancing Action Generation via Pose-Conditioned Anchor Attention
Ziwen Li, Xin Wang, Hanlue Zhang +8
The Vision-Language-Action (VLA) models have demonstrated remarkable performance on embodied tasks and shown promising potential for real-world applications. However, current VLAs…
REMONI: An Autonomous System Integrating Wearables and Multimodal Large Language Models for Enhanced Remote Health Monitoring
Thanh Cong Ho, Farah Kharrat, Abderrazek Abid +1
With the widespread adoption of wearable devices in our daily lives, the demand and appeal for remote patient monitoring have significantly increased. Most research in this field h…
Vision Language Models for Dynamic Human Activity Recognition in Healthcare Settings
Abderrazek Abid, Thanh-Cong Ho, Fakhri Karray
As generative AI continues to evolve, Vision Language Models (VLMs) have emerged as promising tools in various healthcare applications. One area that remains relatively underexplor…
FaceAnonyMixer: Cancelable Faces via Identity Consistent Latent Space Mixing
Mohammed Talha Alam, Fahad Shamshad, Fakhri Karray +1
Advancements in face recognition (FR) technologies have amplified privacy concerns, necessitating methods that protect identity while maintaining recognition utility. Existing face…