16 papers
Towards Robustness against Typographic Attack with Training-free Concept Localization
Bohan Liu, Wenqian Ye, Guangzhi Xiong +3
Models trained via Contrastive Language-Image Pretraining (CLIP) serve as the foundational vision encoders for most modern Large Vision Language Models (LVLMs). Despite their wides…
Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning
Sanchit Sinha, Guangzhi Xiong, Bohan Liu +2
The effectiveness of Chain-of-Thought (CoT) prompting in Multimodal Large Language Models (MLLMs) remains uncertain: across several visual reasoning benchmarks, CoT prompting often…
Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models
Guangzhi Xiong, Qiao Jin, Sanchit Sinha +2
Large Vision Language Models (LVLMs) show promise in medical applications, but their inability to faithfully ground responses in visual evidence raises serious concerns about clini…
Retrieving Counterfactuals Improves Visual In-Context Learning
Guangzhi Xiong, Sanchit Sinha, Zhenghao He +1
Vision-language models (VLMs) have achieved impressive performance across a wide range of multimodal reasoning tasks, but they often struggle to disentangle fine-grained visual att…
Neural Additive Experts: Context-Gated Experts for Controllable Model Additivity
Guangzhi Xiong, Sanchit Sinha, Aidong Zhang
The trade-off between interpretability and accuracy remains a core challenge in machine learning. Standard Generalized Additive Models (GAMs) offer clear feature attributions but a…
Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
Guangzhi Xiong, Zhenghao He, Bohan Liu +2
Retrieval-Augmented Generation (RAG) improves the factuality of large language models (LLMs) by grounding outputs in retrieved evidence, but faithfulness failures, where generation…