2 papers
cs.LG2026
ClawArena: Benchmarking AI Agents in Evolving Information Environments
Haonian Ji, Kaiwen Xiong, Siwei Han +9
AI agents deployed as persistent assistants must maintain correct beliefs as their information environment evolves. In practice, evidence is scattered across heterogeneous sources…
cs.CL2023
CEScore: Simple and Efficient Confidence Estimation Model for Evaluating Split and Rephrase
AlMotasem Bellah Al Ajlouni, Jinlong Li
The split and rephrase (SR) task aims to divide a long, complex sentence into a set of shorter, simpler sentences that convey the same meaning. This challenging problem in NLP has…