1 citations · 2 across the 9 of their papers we have counts for
7 papers · 1 filter
Evaluation Framework for AI Systems in "the Wild"
Sarah Jabbour, Trenton Chang, Anindya Das Antar +13
Generative AI (GenAI) models have become vital across industries, yet current evaluation methods have not adapted to their widespread use. Traditional evaluations often rely on ben…
Towards Dog Bark Decoding: Leveraging Human Speech Processing for Automated Bark Classification
Artem Abzaliev, Humberto Pérez Espinosa, Rada Mihalcea
Similar to humans, animals make extensive use of verbal and non-verbal forms of communication, including a large range of audio signals. In this paper, we address dog vocalizations…
: "This is My SQL, Are You With Me?" A Consensus-Based Multi-Agent System for Text-to-SQL Tasks
Hanchen Xia, Feng Jiang, Naihao Deng +4
Large Language Models (LLMs) have demonstrated strong performance on various tasks. To unleash their power on the Text-to-SQL task, we propose (Review-Rebuttal-Revision), a c…
EmoBench: Evaluating the Emotional Intelligence of Large Language Models
Sahand Sabour, Siyang Liu, Zheyuan Zhang +7
Recent advances in Large Language Models (LLMs) have highlighted the need for robust, comprehensive, and challenging benchmarks. Yet, research on evaluating their Emotional Intelli…
You Are What You Annotate: Towards Better Models through Annotator Representations
Naihao Deng, Xinliang Frederick Zhang, Siyang Liu +3
Annotator disagreement is ubiquitous in natural language processing (NLP) tasks. There are multiple reasons for such disagreements, including the subjectivity of the task, difficul…
STaCK: Sentence Ordering with Temporal Commonsense Knowledge
Deepanway Ghosal, Navonil Majumder, Rada Mihalcea +1
Sentence order prediction is the task of finding the correct order of sentences in a randomly ordered document. Correctly ordering the sentences requires an understanding of cohere…