5 papers
Loki: Representation over Architecture for Diffusion-Based Portrait Animation
Pouyan Navard, Sernam Lim
Portrait animation transfers a driver clip's facial expression and head pose onto a single reference image while preserving the reference's identity. State-of-the-art diffusion sys…
DiReCT: Disentangled Regularization of Contrastive Trajectories for Physics-Refined Video Generation
Abolfazl Meyarian, Amin Karimi Monsefi, Rajiv Ramnath +1
Flow-matching video generators produce temporally coherent, high-fidelity outputs yet routinely violate elementary physics because their reconstruction objectives penalize per-fram…
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness
Ahmad Mohammadshirazi, Pinaki Prasad Guha Neogi, Ser-Nam Lim +1
Document Visual Question Answering (VQA) demands robust integration of text detection, recognition, and spatial reasoning to interpret complex document layouts. In this work, we in…
DetailCLIP: Detail-Oriented CLIP for Fine-Grained Tasks
Amin Karimi Monsefi, Kishore Prakash Sailaja, Ali Alilooee +2
In this paper, we introduce DetailCLIP: A Detail-Oriented CLIP to address the limitations of contrastive learning-based vision-language models, particularly CLIP, in handling detai…
Frequency-Guided Masking for Enhanced Vision Self-Supervised Learning
Amin Karimi Monsefi, Mengxi Zhou, Nastaran Karimi Monsefi +3
We present a novel frequency-based Self-Supervised Learning (SSL) approach that significantly enhances its efficacy for pre-training. Prior work in this direction masks out pre-def…