4 papers
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
Tyler A. Chang, Catherine Arnett, Abdelrahman Sadallah +377
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we pre…
Testing the Limits of Truth Directions in LLMs
Angelos Poulis, Mark Crovella, Evimaria Terzi
Large language models (LLMs) have been shown to encode truth of statements in their activation space along a linear truth direction. Previous studies have argued that these directi…
Transformer-based Language Models for Reasoning in the Description Logic ALCQ
Angelos Poulis, Eleni Tsalapati, Manolis Koubarakis
Recent advancements in transformer-based language models have sparked research into their logical reasoning capabilities. Most of the benchmarks used to evaluate these models are s…
Transformers in the Service of Description Logic-based Contexts
Angelos Poulis, Eleni Tsalapati, Manolis Koubarakis
Recent advancements in transformer-based models have initiated research interests in investigating their ability to learn to perform reasoning tasks. However, most of the contexts…