From the 1 of 6 linked papers with an AI index.
6 papers
Is ChatGPT as reliable as individual reviewers assessing the quality of published journal articles from PDFs or titles and abstracts?
Mike Thelwall, Parveen Ali
The paper compares how well ChatGPT-5.4 can rate the research quality of journal articles against scores given by human reviewers, using both title/abstract and full‑text PDF input…
Can ChatGPT be a good follower of academic paradigms? Research quality evaluations in conflicting areas of sociology
Mike Thelwall, Ralph Schroeder, Meena Dhanda
Purpose: It has become increasingly likely that Large Language Models (LLMs) will be used to score the quality of academic publications to support research assessment goals in the…
Can Small and Reasoning Large Language Models Score Journal Articles for Research Quality and Do Averaging and Few-shot Help?
Mike Thelwall, Ehsan Mohammadi
Previous research has shown that journal article quality ratings from the cloud based Large Language Model (LLM) families ChatGPT and Gemini and the medium sized open weights LLM G…
Large Language Models for Departmental Expert Review Quality Scores
Liv Langfeldt, Dag W. Aksnes, Henrik Karlstrøm +1
Presumably, peer reviewers and Large Language Models (LLMs) do very different things when asked to assess research. Still, recent evidence has shown that LLMs have a moderate abili…
Designing large language model prompts to extract scores from messy text: A shared dataset and challenge
Mike Thelwall
In some areas of computing, natural language processing and information science, progress is made by sharing datasets and challenging the community to design the best algorithm for…
Prompt perturbation and fraction facilitation sometimes strengthen Large Language Model scores
Mike Thelwall
Large Language Models (LLMs) can be tasked with scoring texts according to pre-defined criteria and on a defined scale, but there is no recognised optimal prompting strategy for th…