5 papers
Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs
Yi Tang, Xinyi Shang, Jiacheng Cui +12
Modern vision-language models (VLMs) have significantly improved image generation and editing capabilities, making pixel-level image tampering detection increasingly important yet…
Probing CLIP's Comprehension of 360-Degree Textual and Visual Semantics
Hai Wang, Xiaochen Yang, Mingzhi Dong +1
The dream of instantly creating rich 360-degree panoramic worlds from text is rapidly becoming a reality, yet a crucial gap exists in our ability to reliably evaluate their semanti…
From Masks to Pixels and Meaning: A New Taxonomy, Benchmark, and Metrics for VLM Image Tampering
Xinyi Shang, Yi Tang, Jiacheng Cui +9
Existing tampering detection benchmarks largely rely on object masks, which severely misalign with the true edit signal: many pixels inside a mask are untouched or only trivially m…
A Survey on Text-Driven 360-Degree Panorama Generation
Hai Wang, Xiaoyu Xiang, Weihao Xia +1
The advent of text-driven 360-degree panorama generation, enabling the synthesis of 360-degree panoramic images directly from textual descriptions, marks a transformative advanceme…
360PanT: Training-Free Text-Driven 360-Degree Panorama-to-Panorama Translation
Hai Wang, Jing-Hao Xue
Preserving boundary continuity in the translation of 360-degree panoramas remains a significant challenge for existing text-driven image-to-image translation methods. These methods…