15 papers
StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models
Siyu Xu, Yunke Wang, Zijian Wang +6
Vision-Language-Action (VLA) models can follow instructions and manipulate objects, but their performance often collapses out of distribution (OOD), when the scene, viewpoint, or o…
Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution
Lewei Xu, Yihao Ding, Zihan Xu +5
Multi-page visually-rich document understanding (MP-VRDU) requires managing evidence that is sparse, spread across pages, and often exceeds a model's context window. Prior work has…
Rethinking Visual Autoregressive Sampling with Information-Grounding Guidance
Ky Dan Nguyen, Hoang Lam Tran, Anh-Dung Dinh +4
Autoregressive (AR) models based on next-scale prediction have emerged as a powerful tool for image generation, but they face a critical weakness: information inconsistencies betwe…
Multi-scale Coarse-to-fine Modeling for Test-time Human Motion Control
Nhat Le, Daochang Liu, Anh Nguyen +1
We present MSCoT, a multi-scale, coarse-to-fine model for test-time human motion synthesis and control. Unlike recent approaches that rely on multiple iterative denoising/token-pre…
Implicit Neural Representation-Based Continuous Single Image Super-Resolution: An Empirical Benchmark
Tayyab Nasir, Daochang Liu, Ajmal Mian
Implicit neural representation (INR) has become the standard approach for arbitrary-scale image super-resolution (ASSR). To date, no empirical study has systematically examined the…
NAIMA: Semantics Aware RGB Guided Depth Super-Resolution
Tayyab Nasir, Daochang Liu, Ajmal Mian
Guided depth super-resolution (GDSR) is a multi-modal approach for depth map super-resolution that relies on a low-resolution depth map and a high-resolution RGB image to restore f…