13 papers · 1 filter
Domain-Grounded Candidate Selection for Agentic Image Editing: A Shadow Removal Case
Shilin Hu, Jingyi Xu, Dimitris Samaras +1
Commercial vision-language models are reshaping computer vision, with visual priors broad enough to rival task-specific systems. This raises a natural question: do they reduce the…
OmniGF: A Dual-Branch Vision-Language Framework for Unified Gaze Following
Qiaomu Miao, Haoyu Wu, Jingyi Xu +2
Understanding human gaze behavior is essential for complex scene comprehension and human-computer interaction. Traditional gaze following models are typically restricted to pure sp…
Automated Counting of Stacked Objects in Industrial Inspection
Corentin Dumery, Noa Etté, Aoxiang Fan +4
Visual object counting is a fundamental computer vision task in industrial inspection, where accurate, high-throughput inventory tracking and quality assurance are critical. Moreov…
Personalized Image Descriptions from Attention Sequences
Ruoyu Xue, Hieu Le, Jingyi Xu +5
People can view the same image differently: they focus on different regions, objects, and details in varying orders and describe them in distinct linguistic styles. This leads to s…
Cast and Attached Shadow Detection via Iterative Light and Geometry Reasoning
Shilin Hu, Jingyi Xu, Sagnik Das +2
Shadows encode rich information about scene geometry and illumination, yet existing methods either predict a unified shadow mask or overlook attached shadows entirely. We address t…
Embedding Physical Reasoning into Diffusion-Based Shadow Generation
Shilin Hu, Jingyi Xu, Akshat Dave +2
Generating realistic shadows for inserted objects requires reasoning about scene geometry and illumination. However, most existing methods operate purely in image space, leaving th…