4 papers
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
Shuzheng Gao, Eric John Li, Man Ho Lam +5
Large foundation models are fundamentally transforming the software engineering landscape, demonstrating exceptional capabilities across diverse tasks such as code generation, debu…
How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments
Jen-tse Huang, Eric John Li, Man Ho Lam +7
Decision-making is a complex process requiring diverse abilities, making it an excellent framework for evaluating Large Language Models (LLMs). Researchers have examined LLMs' deci…
Emotionally Numb or Empathetic? Evaluating How LLMs Feel Using EmotionBench
Jen-tse Huang, Man Ho Lam, Eric John Li +5
Evaluating Large Language Models' (LLMs) anthropomorphic capabilities has become increasingly important in contemporary discourse. Utilizing the emotion appraisal theory from psych…
Revisiting the Reliability of Psychological Scales on Large Language Models
Jen-tse Huang, Wenxiang Jiao, Man Ho Lam +3
Recent research has focused on examining Large Language Models' (LLMs) characteristics from a psychological standpoint, acknowledging the necessity of understanding their behaviora…