2 citations · 3 across the 3 of their papers we have counts for
7 papers · 1 filter
Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests
Patrick Rim, Tom Long, Ekta Prashnani +6
Multimodal large language models (MLLMs) excel at visual interpretation but fail on spatial reasoning tasks that humans solve reliably. Existing benchmarks evaluate these models as…
Unmasking Puppeteers: Leveraging Biometric Leakage to Expose Impersonation in AI-Based Videoconferencing
Danial Samadi Vahdati, Tai Duc Nguyen, Ekta Prashnani +4
AI-based talking-head videoconferencing systems reduce bandwidth by sending a compact pose-expression latent and re-synthesizing RGB at the receiver, but this latent can be puppete…
Seeing What Matters: Generalizable AI-generated Video Detection with Forensic-Oriented Augmentation
Riccardo Corvi, Davide Cozzolino, Ekta Prashnani +3
Synthetic video generation is progressing very rapidly. The latest models can produce very realistic high-resolution videos that are virtually indistinguishable from real ones. Alt…
Avatar Fingerprinting for Authorized Use of Synthetic Talking-Head Videos
Ekta Prashnani, Koki Nagano, Shalini De Mello +2
Modern avatar generators allow anyone to synthesize photorealistic real-time talking avatars, ushering in a new era of avatar-based human communication, such as with immersive AR/V…
Generalizable Deepfake Detection with Phase-Based Motion Analysis
Ekta Prashnani, Michael Goebel, B. S. Manjunath
We propose PhaseForensics, a DeepFake (DF) video detection method that leverages a phase-based motion representation of facial temporal dynamics. Existing methods relying on tempor…
LOCL: Learning Object-Attribute Composition using Localization
Satish Kumar, ASM Iftekhar, Ekta Prashnani +1
This paper describes LOCL (Learning Object Attribute Composition using Localization) that generalizes composition zero shot learning to objects in cluttered and more realistic sett…