6 papers · 1 filter
InstructMixup: Instruction-Guided Salient Patch Editing for Robust Data Augmentation
Khawar Islam, Arif Mahmood, Xin Jin +1
In image and video technologies, data augmentation is widely used to improve the generalization of deep visual models, and mixup-based strategies that interpolate between samples h…
-FracMix: Label-Preserving Self-Saliency Mixup Augmentation
Khawar Islam, Arif Mahmood, Xin Jin +1
Data augmentation is known to improve generalization of deep visual models. Recent methods favor mixup strategies that generate interpolated samples to improve model performance. H…
Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity
Heethanjan Kanagalingam, Thenukan Pathmanathan, Mokeeshan Vathanakumar +3
Vision Language Models have achieved near-human performance on single-document Visual Question Answering, yet their effectiveness degrades significantly when retrieving information…
Symmetry-Aware 9D Pose Estimation with Sim(3)-Consistent Feature and Spherical Inception Convolution
Panfei Cheng, Hongshan Yu, Wenrui Chen +3
Object pose estimation is a fundamental problem for an agent system to perceive or manipulate objects in images or videos. However, current instance-level methods struggle with gen…
Latent Video Prediction Learns Better World Models
Ali J Alrasheed, Aryan Yazdan Parast, Basim Azam +2
Self-supervised video models are increasingly framed as world models, yet their evaluation remains largely confined to a single top-1 accuracy score on clean benchmarks. This leave…
Mitigating Memorization in Text-to-Image Diffusion via Region-Aware Prompt Augmentation and Multimodal Copy Detection
Yunzhuo Chen, Jordan Vice, Naveed Akhtar +2
State-of-the-art text-to-image diffusion models can produce impressive visuals but may memorize and reproduce training images, creating copyright and privacy risks. Existing prompt…