4 papers
GENIE: A Fine-Grained Measure for Novelty
Ramya Namuduri, Manya Wadhwa, Anshun Asher Zheng +2
Large Language Models have consistently demonstrated a lack of creativity and diversity across tasks. Prior work has focused on addressing whether models are capable of generating…
HERO'S JOURNEY: Testing Complex Rule Induction with Text Games
Anshun Asher Zheng, Kanishka Misra, David I. Beaver +1
We introduce HERO'S JOURNEY, a benchmark for rule induction in goal-directed episodic tasks, where agents must infer hidden rules from demonstrations and act on them through multi-…
Strategic Dialogue Assessment: The Crooked Path to Innocence
Anshun Asher Zheng, Junyi Jessy Li, David I. Beaver
Language is often used strategically, particularly in high-stakes, adversarial settings, yet most work on pragmatics and LLMs centers on cooperativity. This leaves a gap in the sys…
QUDsim: Quantifying Discourse Similarities in LLM-Generated Text
Ramya Namuduri, Yating Wu, Anshun Asher Zheng +3
As large language models become increasingly capable at various writing tasks, their weakness at generating unique and creative content becomes a major liability. Although LLMs hav…