6 citations · 6 across the 6 of their papers we have counts for
9 papers · 1 filter
MI-Pruner: Crossmodal Mutual Information-guided Token Pruner for Efficient MLLMs
Jiameng Li, Aleksei Tiulpin, Matthew B. Blaschko
For multimodal large language models (MLLMs), visual information is relatively sparse compared with text. As a result, research on visual pruning emerges for efficient inference. C…
Spectrum Matching: a Unified Perspective for Superior Diffusability in Latent Diffusion
Mang Ning, Mingxiao Li, Le Zhang +4
In this paper, we study the diffusability (learnability) of variational autoencoders (VAE) in latent diffusion. First, we show that pixel-space diffusion trained with an MSE object…
When normalization hallucinates: unseen risks in AI-powered whole slide image processing
Karel Moens, Matthew B. Blaschko, Tinne Tuytelaars +3
Whole slide image (WSI) normalization remains a vital preprocessing step in computational pathology. Increasingly driven by deep learning, these models learn to approximate data di…
CLASH: A Benchmark for Cross-Modal Contradiction Detection
Teodora Popordanoska, Jiameng Li, Matthew B. Blaschko
Contradictory multimodal inputs are common in real-world settings, yet existing benchmarks typically assume input consistency and fail to evaluate cross-modal contradiction detecti…
SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model
Dongli Xu, Aleksei Tiulpin, Matthew B. Blaschko
Autoregressive (AR) models have emerged as powerful tools for image generation by modeling images as sequences of discrete tokens. While Classifier-Free Guidance (CFG) has been ado…
Jigsaw-R1: A Study of Rule-based Visual Reinforcement Learning with Jigsaw Puzzles
Zifu Wang, Junyi Zhu, Bo Tang +4
The application of rule-based reinforcement learning (RL) to multimodal large language models (MLLMs) introduces unique challenges and potential deviations from findings in text-on…