145 citations · 190 across the 16 of their papers we have counts for
4 papers · 1 filter
Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards
Minji Jung, Minjae Lee, Yejin Kim +2
LLM leaderboards are widely used to compare models and guide deployment decisions. However, leaderboard rankings are shaped by evaluation priorities set by benchmark designers, rat…
COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs
Dasol Choi, DongGeon Lee, Brigitta Jesica Kartono +6
As large language models are deployed in high-stakes enterprise applications, from healthcare to finance, ensuring adherence to organization-specific policies has become essential.…
SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment
Wonje Jeung, Sangyeon Yoon, Minsuk Kahng +1
Large Reasoning Models (LRMs) have become powerful tools for complex problem solving, but their structured reasoning pathways can lead to unsafe outputs when exposed to harmful pro…
Identifying Reasoning Flaws in Planning-Based RL Using Tree Explanations
Kin-Ho Lam, Zhengxian Lin, Jed Irvine +5
Enabling humans to identify potential flaws in an agent's decision making is an important Explainable AI application. We consider identifying such flaws in a planning-based deep re…