activity
20242026
most citedHumanity's Last Exam

18 citations · 21 across the 9 of their papers we have counts for

collaborators
Showing cs.CYShow all

5 papers · 1 filter

cs.CY20261 cited

Measuring and mitigating overreliance to build human-compatible AI

Lujain Ibrahim, Katherine M. Collins, Sunnie S. Y. Kim +14

Large language models (LLMs) distinguish themselves from previous technologies by functioning as collaborative ``thought partners,'' capable of engaging more fluidly in natural lan…

cs.CY20261 cited

On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective

Yue Huang, Chujie Gao, Siyuan Wu +63

Generative Foundation Models (GenFMs) have emerged as transformative tools. However, their widespread adoption raises critical concerns regarding trustworthiness across dimensions.…

cs.CY2025

Prioritization First, Principles Second: An Adaptive Interpretation of Helpful, Honest, and Harmless Principles

Yue Huang, Chujie Gao, Yujun Zhou +5

The Helpful, Honest, and Harmless (HHH) principle is a foundational framework for aligning AI systems with human values. However, existing interpretations of the HHH principle ofte…

cs.CY2024

Human vs. Machine: Behavioral Differences Between Expert Humans and Language Models in Wargame Simulations

Max Lamparth, Anthony Corso, Jacob Ganz +3

To some, the advent of artificial intelligence (AI) promises better decision-making and increased military effectiveness while reducing the influence of human error and emotions. H…

cs.CY2024

Risks from Language Models for Automated Mental Healthcare: Ethics and Structure for Implementation

Declan Grabb, Max Lamparth, Nina Vasan

Amidst the growing interest in developing task-autonomous AI for automated mental health care, this paper addresses the ethical and practical challenges associated with the issue a…