8 papers
InstructMixup: Instruction-Guided Salient Patch Editing for Robust Data Augmentation
Khawar Islam, Arif Mahmood, Xin Jin +1
In image and video technologies, data augmentation is widely used to improve the generalization of deep visual models, and mixup-based strategies that interpolate between samples h…
Reasoning as a Double-Edged Sword: Architecture and Cross-Stage Robustness in Vision-Language-Action Models
Tuan Duong Trinh, Naveed Akhtar, Basim Azam
Does adding a reasoning step make a Vision-Language-Action (VLA) model more robust to perturbation? Intuitively, a policy that reasons before acting should absorb a perturbed input…
Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity
Heethanjan Kanagalingam, Thenukan Pathmanathan, Mokeeshan Vathanakumar +3
Vision Language Models have achieved near-human performance on single-document Visual Question Answering, yet their effectiveness degrades significantly when retrieving information…
-FracMix: Label-Preserving Self-Saliency Mixup Augmentation
Khawar Islam, Arif Mahmood, Xin Jin +1
Data augmentation is known to improve generalization of deep visual models. Recent methods favor mixup strategies that generate interpolated samples to improve model performance. H…
Symmetry-Aware 9D Pose Estimation with Sim(3)-Consistent Feature and Spherical Inception Convolution
Panfei Cheng, Hongshan Yu, Wenrui Chen +3
Object pose estimation is a fundamental problem for an agent system to perceive or manipulate objects in images or videos. However, current instance-level methods struggle with gen…
Latent Video Prediction Learns Better World Models
Ali J Alrasheed, Aryan Yazdan Parast, Basim Azam +2
Self-supervised video models are increasingly framed as world models, yet their evaluation remains largely confined to a single top-1 accuracy score on clean benchmarks. This leave…