4 papers
Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning
Isabella Luong, Joyee Chen, Arturs Kanepajs +5
Evaluating animal welfare reasoning in LLMs remains an open challenge despite rapid deployment in consumer and professional contexts where welfare considerations appear implicitly…
LCFO: Long Context and Long Form Output Dataset and Benchmarking
Marta R. Costa-jussÃ, Pierre Andrews, Mariano Coria Meglioli +10
This paper presents the Long Context and Form Output (LCFO) benchmark, a novel evaluation framework for assessing gradual summarization and summary expansion capabilities across di…
Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset
Vasu Agrawal, Akinniyi Akinyemi, Kathryn Alvero +81
Human communication involves a complex interplay of verbal and nonverbal signals, essential for conveying meaning and achieving interpersonal goals. To develop socially intelligent…
Y-NQ: English-Yorùbá Evaluation dataset for Open-Book Reading Comprehension and Text Generation
Marta R. Costa-jussÃ, Joy Chen, Ifeoluwanimi Adebara +3
The purpose of this work is to share an English-Yorùbá evaluation dataset for open-book reading comprehension and text generation to assess the performance of models both in a hi…