5 papers
Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer
Haoyu Wu, Jingyi Xu, Qiaomu Miao +2
Rotary positional embeddings (RoPE) are widely used in diffusion transformers (DiTs) to encode spatial relationships, yet their behavior with mixed-resolution tokens remains undere…
OmniGF: A Dual-Branch Vision-Language Framework for Unified Gaze Following
Qiaomu Miao, Haoyu Wu, Jingyi Xu +2
Understanding human gaze behavior is essential for complex scene comprehension and human-computer interaction. Traditional gaze following models are typically restricted to pure sp…
Multi-view Gaze Target Estimation
Qiaomu Miao, Vivek Raju Golani, Jingyi Xu +3
This paper presents a method that utilizes multiple camera views for the gaze target estimation (GTE) task. The approach integrates information from different camera views to impro…
Study of detecting behavioral signatures within DeepFake videos
Qiaomu Miao, Sinhwa Kang, Stacy Marsella +3
There is strong interest in the generation of synthetic video imagery of people talking for various purposes, including entertainment, communication, training, and advertisement. W…
Diffusion-Refined VQA Annotations for Semi-Supervised Gaze Following
Qiaomu Miao, Alexandros Graikos, Jingwei Zhang +3
Training gaze following models requires a large number of images with gaze target coordinates annotated by human annotators, which is a laborious and inherently ambiguous process.…