4 citations · 10 across the 9 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
Xiangyi Li, Kyoung Whan Choe, Yimin Liu +12
Large language model (LLM) agents are increasingly deployed to automate productivity tasks (e.g., email, scheduling, document management), but evaluating them on live services is r…
cs.AI2023★ 2 cited
CodeTransOcean: A Comprehensive Multilingual Benchmark for Code Translation
Weixiang Yan, Yuchen Tian, Yunzhe Li +2
Recent code translation techniques exploit neural machine translation models to translate source code from one programming language to another to satisfy production compatibility o…