most citedLLM Ethics Benchmark: A Three-Dimensional Assessment System for Evaluating Moral Reasoning in Large Language Models

1 citations · 1 across the 1 of their papers we have counts for

collaborators

5 papers

cs.CY2025

Evaluating LLM Safety Across Child Development Stages: A Simulated Agent Approach

Abhejay Murali, Saleh Afroogh, Kevin Chen +3

Current safety alignment for Large Language Models (LLMs) implicitly optimizes for a "modal adult user," leaving models vulnerable to distributional shifts in user cognition. We pr…

cs.CY20251 cited

LLM Ethics Benchmark: A Three-Dimensional Assessment System for Evaluating Moral Reasoning in Large Language Models

Junfeng Jiao, Saleh Afroogh, Abhejay Murali +3

This study establishes a novel framework for systematically evaluating the moral reasoning capabilities of large language models (LLMs) as they increasingly integrate into critical…

cs.LG2025

Reinforcement Learning for Long-Horizon Interactive LLM Agents

Kevin Chen, Marco Cusumano-Towner, Brody Huval +4

Interactive digital agents (IDAs) leverage APIs of stateful digital environments to perform tasks in response to user requests. While IDAs powered by instruction-tuned large langua…

cs.CL2025

AGGA: A Dataset of Academic Guidelines for Generative AI and Large Language Models

Junfeng Jiao, Saleh Afroogh, Kevin Chen +2

This study introduces AGGA, a dataset comprising 80 academic guidelines for the use of Generative AIs (GAIs) and Large Language Models (LLMs) in academic settings, meticulously col…

cs.CY2025

IGGA: A Dataset of Industrial Guidelines and Policy Statements for Generative AIs

Junfeng Jiao, Saleh Afroogh, Kevin Chen +2

This paper introduces IGGA, a dataset of 160 industry guidelines and policy statements for the use of Generative AIs (GAIs) and Large Language Models (LLMs) in industry and workpla…