6 papers
C-PTQ: Fisher-weighted Channel-wise Sensitivity for Post-training Quantization of MLLMs
Jiameng Li, Han Zhou, Matthew B. Blaschko
Multimodal large language models (MLLMs) require huge memory and computational costs, which limits their practical deployment. Post-training quantization (PTQ) techniques offer an…
EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs
Jiameng Li, Minye Wu, Jiezhang Cao +2
Long-form video understanding remains challenging for Video Large Language Models (VideoLLMs), as the dense frame sampling introduces massive visual tokens while sparse sampling ri…
MI-Pruner: Crossmodal Mutual Information-guided Token Pruner for Efficient MLLMs
Jiameng Li, Aleksei Tiulpin, Matthew B. Blaschko
For multimodal large language models (MLLMs), visual information is relatively sparse compared with text. As a result, research on visual pruning emerges for efficient inference. C…
CARE: Confidence-aware Ratio Estimation for Medical Biomarkers
Jiameng Li, Teodora Popordanoska, Aleksei Tiulpin +3
Ratio-based biomarkers (RBBs), such as the proportion of necrotic tissue within a tumor, are widely used in clinical practice to support diagnosis, prognosis, and treatment plannin…
CLASH: A Benchmark for Cross-Modal Contradiction Detection
Teodora Popordanoska, Jiameng Li, Matthew B. Blaschko
Contradictory multimodal inputs are common in real-world settings, yet existing benchmarks typically assume input consistency and fail to evaluate cross-modal contradiction detecti…
YpathRAG:A Retrieval-Augmented Generation Framework and Benchmark for Pathology
Deshui Yu, Yizhi Wang, Saihui Jin +9
Large language models (LLMs) excel on general tasks yet still hallucinate in high-barrier domains such as pathology. Prior work often relies on domain fine-tuning, which neither ex…