3 papers
cs.CV2026
Channel-wise Vector Quantization
Wei Song, Tianhang Wang, Yitong Chen +5
We present Channel-wise Vector Quantization (CVQ), a novel image tokenization paradigm that replaces patch-wise tokens with channel-wise tokens. Unlike conventional vector quantiza…
cs.CV2025
VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning
Zhaozhi Wang, Tong Zhang, Mingyue Guo +2
Multimodal Large Language Models (MLLMs) have achieved impressive progress in vision-language alignment, yet they remain limited in visual-spatial reasoning. We first identify that…
cs.CV2024
Mixture of Physical Priors Adapter for Parameter-Efficient Fine-Tuning
Zhaozhi Wang, Conghu Li, Qixiang Ye +1
Most parameter-efficient fine-tuning (PEFT) methods rely on low-rank representations to adapt models. However, these approaches often oversimplify representations, particularly whe…