activity
20242026
collaborators

15 papers

cs.RO2026

StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models

Siyu Xu, Yunke Wang, Zijian Wang +6

Vision-Language-Action (VLA) models can follow instructions and manipulate objects, but their performance often collapses out of distribution (OOD), when the scene, viewpoint, or o…

cs.AI2026

Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution

Lewei Xu, Yihao Ding, Zihan Xu +5

Multi-page visually-rich document understanding (MP-VRDU) requires managing evidence that is sparse, spread across pages, and often exceeds a model's context window. Prior work has…

cs.CV2026

Rethinking Visual Autoregressive Sampling with Information-Grounding Guidance

Ky Dan Nguyen, Hoang Lam Tran, Anh-Dung Dinh +4

Autoregressive (AR) models based on next-scale prediction have emerged as a powerful tool for image generation, but they face a critical weakness: information inconsistencies betwe…

cs.CV2026

Multi-scale Coarse-to-fine Modeling for Test-time Human Motion Control

Nhat Le, Daochang Liu, Anh Nguyen +1

We present MSCoT, a multi-scale, coarse-to-fine model for test-time human motion synthesis and control. Unlike recent approaches that rely on multiple iterative denoising/token-pre…

cs.CV2026

Implicit Neural Representation-Based Continuous Single Image Super-Resolution: An Empirical Benchmark

Tayyab Nasir, Daochang Liu, Ajmal Mian

Implicit neural representation (INR) has become the standard approach for arbitrary-scale image super-resolution (ASSR). To date, no empirical study has systematically examined the…

eess.IV2026

NAIMA: Semantics Aware RGB Guided Depth Super-Resolution

Tayyab Nasir, Daochang Liu, Ajmal Mian

Guided depth super-resolution (GDSR) is a multi-modal approach for depth map super-resolution that relies on a low-resolution depth map and a high-resolution RGB image to restore f…