7 papers
Prompt-Guided Image Editing with Masked Logit Nudging in Visual Autoregressive Models
Amir El-Ghoussani, Marc Hölle, Gustavo Carneiro +1
We address the problem of prompt-guided image editing in visual autoregressive models. Given a source image and a target text prompt, we aim to modify the source image according to…
Visual Autoregressive Modelling for Monocular Depth Estimation
Amir El-Ghoussani, André Kaup, Nassir Navab +2
We propose a monocular depth estimation method based on visual autoregressive (VAR) priors, offering an alternative to diffusion-based approaches. Our method adapts a large-scale t…
TransPrune: Token Transition Pruning for Efficient Large Vision-Language Model
Ao Li, Yuxiang Duan, Jinghui Zhang +5
Large Vision-Language Models (LVLMs) have advanced multimodal learning but face high computational costs due to the large number of visual tokens, motivating token pruning to impro…
Rethinking Weight-Averaged Model-merging
Hu Wang, Congbo Ma, Ibrahim Almakky +3
Model merging, particularly through weight averaging, has shown surprising effectiveness in saving computations and improving model performance without any additional training. How…
ItTakesTwo: Leveraging Peer Representations for Semi-supervised LiDAR Semantic Segmentation
Yuyuan Liu, Yuanhong Chen, Hu Wang +3
The costly and time-consuming annotation process to produce large training sets for modelling semantic LiDAR segmentation methods has motivated the development of semi-supervised l…
A Novel Perspective for Multi-modal Multi-label Skin Lesion Classification
Yuan Zhang, Yutong Xie, Hu Wang +3
The efficacy of deep learning-based Computer-Aided Diagnosis (CAD) methods for skin diseases relies on analyzing multiple data modalities (i.e., clinical+dermoscopic images, and pa…