2 papers
cs.CL2024
NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision
Xiang Li, Wenyue Hua, Kaijie Zhu +8
Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in multimodal understanding, yet their reasoning abilities remain underexplored. Existing benchmarks…
cs.AI2023
NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes
Lizhou Fan, Wenyue Hua, Lingyao Li +2
Complex reasoning ability is one of the most important features of current LLMs, which has also been leveraged to play an integral role in complex decision-making tasks. Therefore,…