12 papers · 1 filter
C-PTQ: Fisher-weighted Channel-wise Sensitivity for Post-training Quantization of MLLMs
Jiameng Li, Han Zhou, Matthew B. Blaschko
Multimodal large language models (MLLMs) require huge memory and computational costs, which limits their practical deployment. Post-training quantization (PTQ) techniques offer an…
EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs
Jiameng Li, Minye Wu, Jiezhang Cao +2
Long-form video understanding remains challenging for Video Large Language Models (VideoLLMs), as the dense frame sampling introduces massive visual tokens while sparse sampling ri…
MI-Pruner: Crossmodal Mutual Information-guided Token Pruner for Efficient MLLMs
Jiameng Li, Aleksei Tiulpin, Matthew B. Blaschko
For multimodal large language models (MLLMs), visual information is relatively sparse compared with text. As a result, research on visual pruning emerges for efficient inference. C…
CARE: Confidence-aware Ratio Estimation for Medical Biomarkers
Jiameng Li, Teodora Popordanoska, Aleksei Tiulpin +3
Ratio-based biomarkers (RBBs), such as the proportion of necrotic tissue within a tumor, are widely used in clinical practice to support diagnosis, prognosis, and treatment plannin…
Spectrum Matching: a Unified Perspective for Superior Diffusability in Latent Diffusion
Mang Ning, Mingxiao Li, Le Zhang +4
In this paper, we study the diffusability (learnability) of variational autoencoders (VAE) in latent diffusion. First, we show that pixel-space diffusion trained with an MSE object…
Revisiting Reweighted Risk for Calibration: AURC, Focal, and Inverse Focal Loss
Han Zhou, Sebastian G. Gruber, Teodora Popordanoska +1
Several variants of reweighted risk functionals, such as focal loss, inverse focal loss, and the Area Under the Risk Coverage Curve (AURC), have been proposed for improving model c…