3 papers
cs.CV2025
Visual Autoregressive Modelling for Monocular Depth Estimation
Amir El-Ghoussani, André Kaup, Nassir Navab +2
We propose a monocular depth estimation method based on visual autoregressive (VAR) priors, offering an alternative to diffusion-based approaches. Our method adapts a large-scale t…
cs.CV2025
TransPrune: Token Transition Pruning for Efficient Large Vision-Language Model
Ao Li, Yuxiang Duan, Jinghui Zhang +5
Large Vision-Language Models (LVLMs) have advanced multimodal learning but face high computational costs due to the large number of visual tokens, motivating token pruning to impro…
cs.CV2024
A Novel Perspective for Multi-modal Multi-label Skin Lesion Classification
Yuan Zhang, Yutong Xie, Hu Wang +3
The efficacy of deep learning-based Computer-Aided Diagnosis (CAD) methods for skin diseases relies on analyzing multiple data modalities (i.e., clinical+dermoscopic images, and pa…