Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
PAUSE: Editable Strategy Artifacts for Long-Form Cultural Story Adaptation
Taaha Kazi, Vasu Sharma, Mohammad Saifullah +1
Generative AI systems increasingly mediate cultural adaptation, but their cultural decisions are often hidden inside prompts, transient model plans, or final prose. We study PAUSE…
cs.CL2026
NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama
Logan Mann, Abdur Rahman, Mohammad Saifullah +2
Long-form serialized audio drama, with arcs that run for 200 to 800 episodes, is a major creative medium and a setting where frontier large language models (LLMs) fail. We benchmar…
cs.CL2024
Large Language Models as User-Agents for Evaluating Task-Oriented-Dialogue Systems
Taaha Kazi, Ruiliang Lyu, Sizhe Zhou +2
Traditionally, offline datasets have been used to evaluate task-oriented dialogue (TOD) models. These datasets lack context awareness, making them suboptimal benchmarks for convers…