3 papers
cs.CL2026
SagaQA: A Multi-hop Reasoning Benchmark for Long-form Narrative Understanding in TV Series
Galann Pennec, Zhengyuan Liu, Nicholas Asher +2
We introduce SagaQA, a long-form video benchmark for multi-hop reasoning over full-length TV series. Existing video reasoning benchmarks often emphasize local understanding of adja…
cs.HC2026
iRULER: Intelligible Rubric-Based User-Defined LLM Evaluation for Revision
Jingwen Bai, Wei Soon Cheong, Philippe Muller +1
Large Language Models (LLMs) have become indispensable for evaluating writing. However, text feedback they provide is often unintelligible, generic, and not specific to user criter…
cs.CL2025
AutoPsyC: Automatic Recognition of Psychodynamic Conflicts from Semi-structured Interviews with Large Language Models
Sayed Muddashir Hossain, Simon Ostermann, Patrick Gebhard +3
Psychodynamic conflicts are persistent, often unconscious themes that shape a person's behaviour and experiences. Accurate diagnosis of psychodynamic conflicts is crucial for effec…