4 citations · 9 across the 7 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
SpecEval: Evaluating Model Adherence to Behavior Specifications
Ahmed Ahmed, Kevin Klyman, Yi Zeng +2
Companies that develop foundation models publish behavioral guidelines they pledge their models will follow, but it remains unclear if models actually do so. While providers such a…
cs.CL2024★ 4 cited
LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output
Elise Karinshak, Amanda Hu, Kewen Kong +4
Immense effort has been dedicated to minimizing the presence of harmful or biased generative content and better aligning AI output to human intention; however, research investigati…