activity
20212026
most citedEvalCrafter: Benchmarking and Evaluating Large Video Generation Models

5 citations · 20 across the 52 of their papers we have counts for

collaborators
Showing cs.CLShow all

41 papers · 1 filter

cs.CL2026

FPCO-Dialog: A Multi-Turn False-Premise Benchmark for Correction and Cooperation in Vision-Language Models

Jiayuan Ma, Yuqi Lu, Weiyang Guo +5

Vision-language models (VLMs) are increasingly deployed in multi-turn settings where users may describe visual content with incorrect assumptions. Yet existing evaluations rarely i…

cs.CL2026

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives

Yingpeng Ma, Jianhao Yan, Bei Shi +6

The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. However, existing research has larg…

cs.CL2026

CRISP: Critical Step Perception for Training Efficient Deep Search Agents

Haosi Mo, Zihao Yan, Ruiqing Zhang +4

Large language models (LLMs) are increasingly extended into deep search agents that solve complex questions through multi-step interaction with external search and browsing tools.…

cs.CL2026

SuCo: Sufficiency-guided Continuous Adaptive Reasoning

Jiahao Wang, Bingyu Liang, Chenhao Hu +5

Despite remarkable performance on complex tasks, Large Reasoning Models (LRMs) often generate excessively long Chain-of-Thoughts (CoT), inflating computational costs even for simpl…

cs.CL2026

Loong: A Human-Like Long Document Translation Agent with Observe-and-Act Adaptive Context Selection

Yutong Wang, Xuebo Liu, Derek F. Wong +5

Document-level translation remains one of the most challenging tasks for large language models, which are constrained by limited context windows that impede global cohesion, while…

cs.CL2026

CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewers

Hexuan Deng, Xiaopeng Ke, Yichen Li +6

Despite the rapid development of AI reviewers, evaluating such systems remains challenging: metrics favor overlap with human reviews over correctness. However, since human reviews…