5 citations · 20 across the 52 of their papers we have counts for
41 papers · 1 filter
FPCO-Dialog: A Multi-Turn False-Premise Benchmark for Correction and Cooperation in Vision-Language Models
Jiayuan Ma, Yuqi Lu, Weiyang Guo +5
Vision-language models (VLMs) are increasingly deployed in multi-turn settings where users may describe visual content with incorrect assumptions. Yet existing evaluations rarely i…
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
Yingpeng Ma, Jianhao Yan, Bei Shi +6
The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. However, existing research has larg…
CRISP: Critical Step Perception for Training Efficient Deep Search Agents
Haosi Mo, Zihao Yan, Ruiqing Zhang +4
Large language models (LLMs) are increasingly extended into deep search agents that solve complex questions through multi-step interaction with external search and browsing tools.…
SuCo: Sufficiency-guided Continuous Adaptive Reasoning
Jiahao Wang, Bingyu Liang, Chenhao Hu +5
Despite remarkable performance on complex tasks, Large Reasoning Models (LRMs) often generate excessively long Chain-of-Thoughts (CoT), inflating computational costs even for simpl…
Loong: A Human-Like Long Document Translation Agent with Observe-and-Act Adaptive Context Selection
Yutong Wang, Xuebo Liu, Derek F. Wong +5
Document-level translation remains one of the most challenging tasks for large language models, which are constrained by limited context windows that impede global cohesion, while…
CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewers
Hexuan Deng, Xiaopeng Ke, Yichen Li +6
Despite the rapid development of AI reviewers, evaluating such systems remains challenging: metrics favor overlap with human reviews over correctness. However, since human reviews…