15 papers
Domain-Grounded Candidate Selection for Agentic Image Editing: A Shadow Removal Case
Shilin Hu, Jingyi Xu, Dimitris Samaras +1
Commercial vision-language models are reshaping computer vision, with visual priors broad enough to rival task-specific systems. This raises a natural question: do they reduce the…
Cast and Attached Shadow Detection via Iterative Light and Geometry Reasoning
Shilin Hu, Jingyi Xu, Sagnik Das +2
Shadows encode rich information about scene geometry and illumination, yet existing methods either predict a unified shadow mask or overlook attached shadows entirely. We address t…
Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer
Haoyu Wu, Jingyi Xu, Qiaomu Miao +2
Rotary positional embeddings (RoPE) are widely used in diffusion transformers (DiTs) to encode spatial relationships, yet their behavior with mixed-resolution tokens remains undere…
OmniGF: A Dual-Branch Vision-Language Framework for Unified Gaze Following
Qiaomu Miao, Haoyu Wu, Jingyi Xu +2
Understanding human gaze behavior is essential for complex scene comprehension and human-computer interaction. Traditional gaze following models are typically restricted to pure sp…
LoRIF: Low-Rank Influence Functions for Scalable Training Data Attribution
Shuangqi Li, Hieu Le, Jingyi Xu +1
Training data attribution (TDA) identifies which training examples most influenced a model's prediction. Influence function methods are a theoretically grounded family of TDA metho…
Embedding Physical Reasoning into Diffusion-Based Shadow Generation
Shilin Hu, Jingyi Xu, Akshat Dave +2
Generating realistic shadows for inserted objects requires reasoning about scene geometry and illumination. However, most existing methods operate purely in image space, leaving th…