collaborators

5 papers

cs.AI2025

Beyond Accuracy: A Geometric Stability Analysis of Large Language Models in Chess Evaluation

Xidan Song, Weiqi Wang, Ruifeng Cao +1

The evaluation of Large Language Models (LLMs) in complex reasoning domains typically relies on performance alignment with ground-truth oracles. In the domain of chess, this standa…

cs.CV2025

Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing

Weixing Wang, Zifeng Ding, Jindong Gu +4

Large Vision-Language Models (LVLMs) with discrete image tokenizers unify multimodal representations by encoding visual inputs into a finite set of tokens. Despite their effectiven…

cs.CL2025

AVerImaTeC: A Dataset for Automatic Verification of Image-Text Claims with Evidence from the Web

Rui Cao, Zifeng Ding, Zhijiang Guo +2

Textual claims are often accompanied by images to enhance their credibility and spread on social media, but this also raises concerns about the spread of misinformation. Existing d…

cs.CL2025

Evaluating LLMs' Assessment of Mixed-Context Hallucination Through the Lens of Summarization

Siya Qi, Rui Cao, Yulan He +1

With the rapid development of large language models (LLMs), LLM-as-a-judge has emerged as a widely adopted approach for text quality evaluation, including hallucination evaluation.…

cs.CV2024

Decompose and Leverage Preferences from Expert Models for Improving Trustworthiness of MLLMs

Rui Cao, Yuming Jiang, Michael Schlichtkrull +1

Multimodal Large Language Models (MLLMs) can enhance trustworthiness by aligning with human preferences. As human preference labeling is laborious, recent works employ evaluation m…