computer-use benchmarking 1interactive agents 1long-horizon tasks 1safety auditing 1tool-use evaluation 1
From the 1 of 6 linked papers with an AI index.
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning
Mengzhao Jia, Zhihan Zhang, Ignacio Cases +3
Multimodal large language models (MLLMs) have rapidly advanced from perception tasks to complex multi-step reasoning, yet reinforcement learning with verifiable rewards (RLVR) ofte…
cs.CL2025
Presenting a Paper is an Art: Self-Improvement Aesthetic Agents for Academic Presentations
Chengzhi Liu, Yuzhe Yang, Kaiwen Zhou +5
The promotion of academic papers has become an important means of enhancing research visibility. However, existing automated methods struggle limited storytelling, insufficient aes…
cs.CL2025
CiteEval: Principle-Driven Citation Evaluation for Source Attribution
Yumo Xu, Peng Qi, Jifan Chen +7
Citation quality is crucial in information-seeking systems, directly influencing trust and the effectiveness of information access. Current evaluation frameworks, both human and au…