7 papers
STAGE: A Full-Screenplay Benchmark for Reasoning over Evolving Stories
Qiuyu Tian, Zequn Liu, Yiding Li +10
Movie screenplays are a demanding testbed for long-form narrative understanding, as characters' goals, beliefs, knowledge, and relationships evolve continuously across scenes. Howe…
VeriGeo: Controllable Geometry Question Generation with Numerical and Analytical Verification
Xiaoxian Duan, Zequn Liu, Yingce Xia
Geometry problem generation is useful for AI-assisted education and multimodal mathematical reasoning, but reliable synthesis remains difficult because the problem statement, diagr…
Narrative Knowledge Weaver: Narrative-Centric Retrieval-Augmented Reasoning for Long-Form Text Understanding
Qiuyu Tian, Fengyi Chen, Yiding Li +9
Long-form narrative QA requires reasoning over evolving story worlds rather than isolated passages: answers may depend on earlier goals, changing character states, social relations…
ForeSci: Evaluating LLM Agents for Forward-Looking AI Research Judgment
Qiuyu Tian, Haojie Yin, Yingce Xia +2
AI research often requires decisions before future evidence exists: which bottleneck to attack, which direction to pursue, or where a project should be positioned. We introduce For…
CoPA: Benchmarking Personalized Question Answering with Data-Informed Cognitive Factors
Hang Su, Zequn Liu, Chen Hu +3
While LLMs have demonstrated remarkable potential in Question Answering (QA), evaluating personalization remains a critical bottleneck. Existing paradigms predominantly rely on lex…
Deciphering Scientific Reasoning Steps from Outcome Data for Molecule Optimization
Zequn Liu, Kehan Wu, Shufang Xie +5
Emerging reasoning models hold promise for automating scientific discovery. However, their training is hindered by a critical supervision gap: experimental outcomes are abundant, w…