4 papers
MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction
Ruoyu Wu, Shenfu Xie, Yinqian Sun +2
Interactive clinical agents operate under partial observability, so reliable care depends on reaching the correct diagnosis through evidence-grounded, safe interactions. Yet existi…
SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?
Kean Shi, Zihang Li, Tianyi Ma +13
Computer-Using Agents (CUAs) are rapidly extending large language models (LLMs) beyond text-based reasoning toward action execution in more complex environments, such as web browse…
RoadmapBench: Evaluating Long-Horizon Agentic Software Development Across Version Upgrades
Xinbo Xu, Ruihan Yang, Haiyang Shen +13
Coding agents are increasingly deployed in real software development, where a single version iteration requires months of coordinated work across many files. However, most existing…
Step-wise Rubric Rewards for LLM Reasoning
Weichu Xie, Haozhe Zhao, Wenpu Liu +15
Reinforcement Learning with Verifiable Rewards (RLVR) is widely used to improve reasoning in large language models, but rewards only final-answer correctness with no supervision ov…