3 papers
cs.AI2026
WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games
Wenyu Zhang, Guoliang You, Tianlun +8
Coding agents are increasingly used as application builders, yet many evaluations still focus on source code, repository-level tests, or intermediate traces rather than the deliver…
cs.RO2026
Towards Generalizable Robotic Data Flywheel: High-Dimensional Factorization and Composition
Yuyang Xiao, Yifei Zhou, Haoran Wang +2
The lack of sufficiently diverse data, coupled with limited data efficiency, remains a major bottleneck for generalist robotic models, yet systematic strategies for collecting and…
cs.CL2024
MIMIR: A Streamlined Platform for Personalized Agent Tuning in Domain Expertise
Chunyuan Deng, Xiangru Tang, Yilun Zhao +5
Recently, large language models (LLMs) have evolved into interactive agents, proficient in planning, tool use, and task execution across a wide variety of tasks. However, without s…