1 citations · 1 across the 8 of their papers we have counts for
5 papers · 1 filter
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction
Tencent WorkBuddy Bench Team, Siqi Cai, Shaopeng Chen +35
We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model lea…
SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents
Shaofei Cai, Yulei Qin, Haojia Lin +10
Agentic reinforcement learning (RL) holds great promise for the development of autonomous agents under complex GUI tasks, but its scalability remains severely hampered by the verif…
LTD-Bench: Evaluating Large Language Models by Letting Them Draw
Liuhao Lin, Ke Li, Zihan Xu +5
Current evaluation paradigms for large language models (LLMs) represent a critical blind spot in AI research--relying on opaque numerical metrics that conceal fundamental limitatio…
Training-Free Group Relative Policy Optimization
Yuzheng Cai, Siqi Cai, Yuchen Shi +10
Recent advances in Large Language Model (LLM) agents have demonstrated their promising general capabilities. However, their performance in specialized real-world domains often degr…
LUCY: Linguistic Understanding and Control Yielding Early Stage of Her
Heting Gao, Hang Shao, Xiong Wang +12
The film Her features Samantha, a sophisticated AI audio agent who is capable of understanding both linguistic and paralinguistic information in human speech and delivering real-ti…