3 papers
cs.CV2026
Geometric Similarity in VLM Low-Level Vision Representations
Shao-Jun Xia, Huixin Zhang, Zhen Lei +3
Vision-language models (VLMs) have emerged as powerful candidates for universal vision backbones, with representative architectures including autoregressive (AR) models and diffusi…
cs.CV2026
Hidden-Shot: Towards One-Shot Task Generalization for Low-Level Vision Generalist Models
Shao-Jun Xia, Xianzheng Ma, Zichong Meng
Despite the intense engagement surrounding low-level vision generalist models, their effectiveness in zero/few-shot scenarios beyond learned tasks remains unverified. The primary c…
cs.CV2025
T2T-VICL: Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs
Shao-Jun Xia, Huixin Zhang, Zhengzhong Tu
Visual in-context learning (VICL) solves visual tasks by conditioning on a few input-output demonstrations without any model training. Recent advances in large vision-language mode…