activity
20232026
most citedEvaluating the Logical Reasoning Ability of ChatGPT and GPT-4

106 citations · 111 across the 5 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

Real Images, Worse Judgments: Evaluating Vision-Language Models on Concreteness and Imagery

Yifan Jiang, Ruoxi Ning, Sheng Yao +1

Visual inputs are often assumed to improve language understanding in multimodal models. We examine this assumption by asking whether vision-language models (VLMs) can distinguish u…

cs.CL2025

From Behavioral Performance to Internal Competence: Interpreting Vision-Language Models with VLM-Lens

Hala Sheta, Eric Huang, Shuyu Wu +9

We introduce VLM-Lens, a toolkit designed to enable systematic benchmarking, analysis, and interpretation of vision-language models (VLMs) by supporting the extraction of intermedi…

cs.CL2024

NovelQA: Benchmarking Question Answering on Documents Exceeding 200K Tokens

Cunxiang Wang, Ruoxi Ning, Boqi Pan +8

Recent advancements in Large Language Models (LLMs) have pushed the boundaries of natural language processing, especially in long-context understanding. However, the evaluation of…

cs.CL2023

GLoRE: Evaluating Logical Reasoning of Large Language Models

Hanmeng liu, Zhiyang Teng, Ruoxi Ning +4

Large language models (LLMs) have shown significant general language understanding abilities. However, there has been a scarcity of attempts to assess the logical reasoning capacit…

cs.CL2023106 cited

Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4

Hanmeng Liu, Ruoxi Ning, Zhiyang Teng +3

Harnessing logical reasoning ability is a comprehensive natural language understanding endeavor. With the release of Generative Pretrained Transformer 4 (GPT-4), highlighted as "ad…