collaborators

8 papers

cs.CV2026

InstructMixup: Instruction-Guided Salient Patch Editing for Robust Data Augmentation

Khawar Islam, Arif Mahmood, Xin Jin +1

In image and video technologies, data augmentation is widely used to improve the generalization of deep visual models, and mixup-based strategies that interpolate between samples h…

cs.RO2026

Reasoning as a Double-Edged Sword: Architecture and Cross-Stage Robustness in Vision-Language-Action Models

Tuan Duong Trinh, Naveed Akhtar, Basim Azam

Does adding a reasoning step make a Vision-Language-Action (VLA) model more robust to perturbation? Intuitively, a policy that reasons before acting should absorb a perturbed input…

cs.CV2026

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity

Heethanjan Kanagalingam, Thenukan Pathmanathan, Mokeeshan Vathanakumar +3

Vision Language Models have achieved near-human performance on single-document Visual Question Answering, yet their effectiveness degrades significantly when retrieving information…

cs.CV2026

-FracMix: Label-Preserving Self-Saliency Mixup Augmentation

Khawar Islam, Arif Mahmood, Xin Jin +1

Data augmentation is known to improve generalization of deep visual models. Recent methods favor mixup strategies that generate interpolated samples to improve model performance. H…

cs.CV2026

Symmetry-Aware 9D Pose Estimation with Sim(3)-Consistent Feature and Spherical Inception Convolution

Panfei Cheng, Hongshan Yu, Wenrui Chen +3

Object pose estimation is a fundamental problem for an agent system to perceive or manipulate objects in images or videos. However, current instance-level methods struggle with gen…

cs.CV2026

Latent Video Prediction Learns Better World Models

Ali J Alrasheed, Aryan Yazdan Parast, Basim Azam +2

Self-supervised video models are increasingly framed as world models, yet their evaluation remains largely confined to a single top-1 accuracy score on clean benchmarks. This leave…