3 papers
cs.CL2026
TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation
Jiahui Zhang, Ziwei Zhang, Yipeng Wang +7
Roleplay evaluation should do more than assign a single score: it should reveal which role requirements were tested, which failed, and which dialogue evidence supports the judgment…
cs.GR2026
Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation
Lanshan He, Haozhou Pang, Qi Gan +12
Cutscenes are carefully choreographed cinematic sequences embedded in video games and interactive media, serving as the primary vehicle for narrative delivery, character developmen…
cs.AI2025
Kimi-Dev: Agentless Training as Skill Prior for SWE-Agents
Zonghan Yang, Shengjie Wang, Kelin Fu +18
Large Language Models (LLMs) are increasingly applied to software engineering (SWE), with SWE-bench as a key benchmark. Solutions are split into SWE-Agent frameworks with multi-tur…