From the 1 of 15 linked papers with an AI index.
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
ProbeLLM: Automating Principled Diagnosis of LLM Failures
Yue Huang, Zhengzhe Jiang, Yuchen Ma +8
Understanding how and why large language models (LLMs) fail is becoming a central challenge as models rapidly evolve and static evaluations fall behind. While automated probing has…
cs.CL2026
NARRA-Gym for Evaluating Interactive Narrative Agents
Yue Huang, Yuchen Ma, Jiayi Ye +14
Interactive narrative tasks require LLMs to sustain a coherent, evolving story while adapting to a user over multiple turns. However, suitable benchmarks for this setting are limit…