1 citations · 1 across the 3 of their papers we have counts for
6 papers · 1 filter
Talking with Tables for Better LLM Factual Data Interactions
Jio Oh, Geon Heo, Seungjun Oh +5
Large Language Models (LLMs) often struggle with requests related to information retrieval and data manipulation that frequently arise in real-world scenarios under multiple condit…
Value Compass Benchmarks: A Platform for Fundamental and Validated Evaluation of LLMs Values
Jing Yao, Xiaoyuan Yi, Shitong Duan +8
As Large Language Models (LLMs) achieve remarkable breakthroughs, aligning their values with humans has become imperative for their responsible development and customized applicati…
General Scales Unlock AI Evaluation with Explanatory and Predictive Power
Lexin Zhou, Lorenzo Pacchiardi, Fernando MartÃnez-Plumed +23
Ensuring safe and effective use of AI requires understanding and anticipating its performance on novel tasks, from advanced scientific challenges to transformed workplace activitie…
PromptBench: A Unified Library for Evaluation of Large Language Models
Kaijie Zhu, Qinlin Zhao, Hao Chen +2
The evaluation of large language models (LLMs) is crucial to assess their performance and mitigate potential security risks. In this paper, we introduce PromptBench, a unified libr…
The Good, The Bad, and Why: Unveiling Emotions in Generative AI
Cheng Li, Jindong Wang, Yixuan Zhang +7
Emotion significantly impacts our daily behaviors and interactions. While recent generative AI models, such as large language models, have shown impressive performance in various t…
DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks
Kaijie Zhu, Jiaao Chen, Jindong Wang +3
Large language models (LLMs) have achieved remarkable performance in various evaluation benchmarks. However, concerns are raised about potential data contamination in their conside…