6 papers
LoC-Path: Learning to Compress for Pathology Multimodal Large Language Models
Qingqiao Hu, Weimin Lyu, Meilong Xu +5
Whole Slide Image (WSI) MLLMs are difficult to build and deploy because gigapixel slides induce thousands of visual tokens, while only a small fraction of regions is diagnostically…
Alignment-Aware Model Adaptation via Feedback-Guided Optimization
Gaurav Bhatt, Aditya Chinchure, Jiawei Zhou +1
Fine-tuning is the primary mechanism for adapting foundation models to downstream tasks; however, standard approaches largely optimize task objectives in isolation and do not accou…
Enhancing Vision-Language Model Reliability with Uncertainty-Guided Dropout Decoding
Yixiong Fang, Ziran Yang, Zhaorun Chen +2
Large vision-language models (LVLMs) excel at multimodal tasks but are prone to misinterpreting visual inputs, often resulting in hallucinations and unreliable outputs. We present…
Features Emerge as Discrete States: The First Application of SAEs to 3D Representations
Albert Miao, Chenliang Zhou, Jiawei Zhou +1
Sparse Autoencoders (SAEs) are a powerful dictionary learning technique for decomposing neural network activations, translating the hidden state into human ideas with high semantic…
Improving Chain-of-Thought Efficiency for Autoregressive Image Generation
Zeqi Gu, Markos Georgopoulos, Xiaoliang Dai +10
Autoregressive multimodal large language models have recently gained popularity for image generation, driven by advances in foundation models. To enhance alignment and detail, newe…
From Behavioral Performance to Internal Competence: Interpreting Vision-Language Models with VLM-Lens
Hala Sheta, Eric Huang, Shuyu Wu +9
We introduce VLM-Lens, a toolkit designed to enable systematic benchmarking, analysis, and interpretation of vision-language models (VLMs) by supporting the extraction of intermedi…