activity
20242026
collaborators

7 papers

cs.CV2026

Prompt-Guided Image Editing with Masked Logit Nudging in Visual Autoregressive Models

Amir El-Ghoussani, Marc Hölle, Gustavo Carneiro +1

We address the problem of prompt-guided image editing in visual autoregressive models. Given a source image and a target text prompt, we aim to modify the source image according to…

cs.CV2025

Visual Autoregressive Modelling for Monocular Depth Estimation

Amir El-Ghoussani, André Kaup, Nassir Navab +2

We propose a monocular depth estimation method based on visual autoregressive (VAR) priors, offering an alternative to diffusion-based approaches. Our method adapts a large-scale t…

cs.CV2025

TransPrune: Token Transition Pruning for Efficient Large Vision-Language Model

Ao Li, Yuxiang Duan, Jinghui Zhang +5

Large Vision-Language Models (LVLMs) have advanced multimodal learning but face high computational costs due to the large number of visual tokens, motivating token pruning to impro…

cs.LG2025

Rethinking Weight-Averaged Model-merging

Hu Wang, Congbo Ma, Ibrahim Almakky +3

Model merging, particularly through weight averaging, has shown surprising effectiveness in saving computations and improving model performance without any additional training. How…

cs.CV2025

ItTakesTwo: Leveraging Peer Representations for Semi-supervised LiDAR Semantic Segmentation

Yuyuan Liu, Yuanhong Chen, Hu Wang +3

The costly and time-consuming annotation process to produce large training sets for modelling semantic LiDAR segmentation methods has motivated the development of semi-supervised l…

cs.CV2024

A Novel Perspective for Multi-modal Multi-label Skin Lesion Classification

Yuan Zhang, Yutong Xie, Hu Wang +3

The efficacy of deep learning-based Computer-Aided Diagnosis (CAD) methods for skin diseases relies on analyzing multiple data modalities (i.e., clinical+dermoscopic images, and pa…