From the 1 of 6 linked papers with an AI index.
6 papers
Benchmarking LLM Competence on Logical Inference over Probability Operators
Nayera Hasan, Jack Greff, Alvin Grissom
The paper presents a benchmark for testing large language models' ability to reason logically about probability expressions in English, and evaluates 29 models, revealing widesprea…
ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation
Ahmed Haj Ahmed, Alvin Grissom
Multilingual reasoning evaluation overwhelmingly relies on translating English benchmarks, a practice that introduces linguistic artifacts and fails to test culturally-grounded rea…
Disentangling Linguistic Relatedness from Task Alignment in Cross-Lingual Transfer
Ahmed Haj Ahmed, Ruochen Zhang, Alvin Grissom
We study cross-lingual transfer by fine-tuning seven large language models (4B--671B parameters) on Arabic and evaluating zero-shot reading comprehension on Semitic languages and n…
An Interdisciplinary Approach to Human-Centered Machine Translation
Marine Carpuat, Omri Asscher, Kalika Bali +17
Machine Translation (MT) tools are widely used today, often in contexts where professional translators are not present. Despite progress in MT technology, a gap persists between sy…
CULEMO: Cultural Lenses on Emotion -- Benchmarking LLMs for Cross-Cultural Emotion Understanding
Tadesse Destaw Belay, Ahmed Haj Ahmed, Alvin Grissom +4
NLP research has increasingly focused on subjective tasks such as emotion analysis. However, existing emotion benchmarks suffer from two major shortcomings: (1) they largely rely o…
Examining Pathological Bias in a Generative Adversarial Network Discriminator: A Case Study on a StyleGAN3 Model
Alvin Grissom, Ryan F. Lei, Matt Gusdorff +3
Generative adversarial networks (GANs) generate photorealistic faces that are often indistinguishable by humans from real faces. While biases in machine learning models are often a…