collaborators

16 papers

cs.CV2026

Adapting Vision Foundation Models with Cascaded Semantics

Xi Xiao, Xingjian Li, Cheng Han +8

Prompt tuning, a leading parameter-efficient adaptation paradigm in NLP, has recently been extended to computer vision. Visual prompt tuning (VPT) adapts pre-trained vision transfo…

cs.LG2026

Less Tokens, Better Forecasts: Sparse Residual Routing for Efficient Weather Prediction

Janet Wang, Yunbei Zhang, Lin Zhao +3

Existing ViT-based weather forecasting models apply uniform computation across all spatial tokens, even though nearby atmospheric grid points often contain similar values and large…

cs.CV2026

Layer-Specific Prompt Fusion Discovery via Differentiable Search in Vision Foundation Models

Xi Xiao, Xingjian Li, Yunbei Zhang +7

Visual prompt tuning has emerged as a parameter-efficient fine-tuning approach for adapting large-scale Vision Transformers (ViTs) to downstream tasks. As its learnable prompts are…

cs.CV2026

FOCUS: Fused Observation of Channels for Unveiling Spectra

Xi Xiao, Aristeidis Tsaris, Anika Tabassum +4

Hyperspectral imaging (HSI) captures hundreds of narrow, contiguous wavelength bands, making it a powerful tool in biology, agriculture, and environmental monitoring. However, inte…

cs.CV2026

Prompt-based Adaptation in Large-scale Vision Models: A Survey

Xi Xiao, Yunbei Zhang, Lin Zhao +12

In computer vision, Visual Prompting (VP) and Visual Prompt Tuning (VPT) have recently emerged as lightweight and effective alternatives to full fine-tuning for adapting large-scal…

cs.CE2025

RoadBench: A Vision-Language Foundation Model and Benchmark for Road Damage Understanding

Xi Xiao, Yunbei Zhang, Janet Wang +9

Accurate road damage detection is crucial for timely infrastructure maintenance and public safety, but existing vision-only datasets and models lack the rich contextual understandi…