3 papers
cs.AI2025
ChemVTS-Bench: Evaluating Visual-Textual-Symbolic Reasoning of Multimodal Large Language Models in Chemistry
Zhiyuan Huang, Baichuan Yang, Zikun He +5
Chemical reasoning inherently integrates visual, textual, and symbolic modalities, yet existing benchmarks rarely capture this complexity, often relying on simple image-text pairs…
cs.CL2025
PiCO: Peer Review in LLMs based on the Consistency Optimization
Kun-Peng Ning, Shuo Yang, Yu-Yang Liu +5
Existing large language models (LLMs) evaluation methods typically focus on testing the performance on some closed-environment and domain-specific benchmarks with human annotations…
cs.CL2024
LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples
Jia-Yu Yao, Kun-Peng Ning, Zhen-Hui Liu +3
Large Language Models (LLMs), including GPT-3.5, LLaMA, and PaLM, seem to be knowledgeable and able to adapt to many tasks. However, we still cannot completely trust their answers,…