6 papers
Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer
Haoyu Wu, Jingyi Xu, Qiaomu Miao +2
Rotary positional embeddings (RoPE) are widely used in diffusion transformers (DiTs) to encode spatial relationships, yet their behavior with mixed-resolution tokens remains undere…
OmniGF: A Dual-Branch Vision-Language Framework for Unified Gaze Following
Qiaomu Miao, Haoyu Wu, Jingyi Xu +2
Understanding human gaze behavior is essential for complex scene comprehension and human-computer interaction. Traditional gaze following models are typically restricted to pure sp…
Learning 3D Reconstruction with Priors in Test Time
Lei Zhou, Haoyu Wu, Akshat Dave +1
We introduce a test-time framework for multiview Transformers (MVTs) that incorporates priors (e.g., camera poses, intrinsics, and depth) to improve 3D tasks without retraining or…
Multi-view Gaze Target Estimation
Qiaomu Miao, Vivek Raju Golani, Jingyi Xu +3
This paper presents a method that utilizes multiple camera views for the gaze target estimation (GTE) task. The approach integrates information from different camera views to impro…
Importance-Based Token Merging for Efficient Image and Video Generation
Haoyu Wu, Jingyi Xu, Hieu Le +1
Token merging can effectively accelerate various vision systems by processing groups of similar tokens only once and sharing the results across them. However, existing token groupi…
MLI-NeRF: Multi-Light Intrinsic-Aware Neural Radiance Fields
Yixiong Yang, Shilin Hu, Haoyu Wu +3
Current methods for extracting intrinsic image components, such as reflectance and shading, primarily rely on statistical priors. These methods focus mainly on simple synthetic sce…