9 papers · 1 filter
MINOS: A Multimodal Evaluation Model for Bidirectional Generation Between Image and Text
Junzhe Zhang, Huixuan Zhang, Xinyu Hu +4
Evaluation is important for multimodal generation tasks, while traditional multimodal evaluation metrics suffer from several limitations. With the rapid progress of MLLMs, there is…
HAD: HAllucination Detection Language Models Based on a Comprehensive Hallucination Taxonomy
Fan Xu, Xinyu Hu, Zhenghan Yu +6
The increasing reliance on natural language generation (NLG) models, particularly large language models, has raised concerns about the reliability and accuracy of their outputs. A…
JointCQ: Improving Factual Hallucination Detection with Joint Claim and Query Generation
Fan Xu, Huixuan Zhang, Zhenliang Zhang +2
Current large language models (LLMs) often suffer from hallucination issues, i,e, generating content that appears factual but is actually unreliable. A typical hallucination detect…
Exploring Causal Effect of Social Bias on Faithfulness Hallucinations in Large Language Models
Zhenliang Zhang, Junzhe Zhang, Xinyu Hu +2
Large language models (LLMs) have achieved remarkable success in various tasks, yet they remain vulnerable to faithfulness hallucinations, where the output does not align with the…
ICR Probe: Tracking Hidden State Dynamics for Reliable Hallucination Detection in LLMs
Zhenliang Zhang, Xinyu Hu, Huixuan Zhang +2
Large language models (LLMs) excel at various natural language processing tasks, but their tendency to generate hallucinations undermines their reliability. Existing hallucination…
Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models
Huixuan Zhang, Xiaojun Wan
Text-to-image models often struggle to generate images that precisely match textual prompts. Prior research has extensively studied the evaluation of image-text alignment in text-t…