most citedHolistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges

13 citations · 33 across the 8 of their papers we have counts for

collaborators

8 papers

cs.LG20246 cited

CAT: Interpretable Concept-based Taylor Additive Models

Viet Duong, Qiong Wu, Zhengyi Zhou +5

As an emerging interpretable technique, Generalized Additive Models (GAMs) adopt neural networks to individually learn non-linear functions for each feature, which are then combine…

cs.CV2024

Generalizing to Unseen Domains in Diabetic Retinopathy with Disentangled Representations

Peng Xia, Ming Hu, Feilong Tang +6

Diabetic Retinopathy (DR), induced by diabetes, poses a significant risk of visual impairment. Accurate and effective grading of DR aids in the treatment of this condition. Yet exi…

cs.CL2024

: Confidence Calibration Model Cascade for Inference-Efficient Cross-Lingual Natural Language Understanding

Taixi Lu, Haoyu Wang, Huajie Shao +2

Cross-lingual natural language understanding (NLU) is a critical task in natural language processing (NLP). Recent advancements have seen multilingual pre-trained language models (…

cs.CL20241 cited

AutoPRM: Automating Procedural Supervision for Multi-Step Reasoning via Controllable Question Decomposition

Zhaorun Chen, Zhuokai Zhao, Zhihong Zhu +4

Recent advancements in large language models (LLMs) have shown promise in multi-step reasoning tasks, yet their reliance on extensive manual labeling to provide procedural feedback…

cs.LG20242 cited

Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Yiyang Zhou, Chenhang Cui, Rafael Rafailov +2

Instruction-following Vision Large Language Models (VLLMs) have achieved significant progress recently on a variety of tasks. These approaches merge strong pre-trained vision model…

cs.CV20241 cited

Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences

Xiyao Wang, Yuhang Zhou, Xiaoyu Liu +9

Multimodal Large Language Models (MLLMs) have demonstrated proficiency in handling a variety of visual-language tasks. However, current MLLM benchmarks are predominantly designed t…