collaborators

15 papers

cs.LG2026

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI

Bohan Lyu, Yucheng Yang, Siqiao Huang +25

Modern AI progress has been driven by ML methods that are generalizable across settings and scalable to larger regimes. As large language models demonstrate advanced capabilities i…

cs.CR2026

Revelio: Cost-Efficient Agentic Memory Safety Vulnerability Detection For Repository-Scale Codebases

Yiwei Hou, Hao Wang, Muxi Lyu +6

Memory safety vulnerabilities remain a significant threat even for projects with extensive fuzzing and manual auditing. Recent results suggest that large language models hold great…

cs.AI2026

ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks

Vincent Siu, Manasi Sharma, Dawn Song +3

Computer use agents are evaluated almost exclusively on atomic desktop tasks, but realistic desktop work requires sustaining state across multiple objectives. We study this gap wit…

cs.AI2026

When Do Intrinsic Rewards Work for Code Reasoning? A Comprehensive Study

Xiaolong Jin, Xuandong Zhao, Wenbo Guo +2

Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in large language model reasoning, but relies on ground-truth supervision that is costly or in…

cs.LG2026

VIMPO: Value-Implicit Policy Optimization for LLMs

Zhewei Kang, Aosong Feng, Sergey Levine +2

Reinforcement learning with verifiable rewards has become a central tool for improving the reasoning ability of large language models, but current methods face a trade-off between…

cs.AI2026

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack

Hao Wang, Hanchen Li, Qiuyang Mang +3

Agent benchmarks have become the de facto measure of frontier AI competence, guiding model selection, investment, and deployment. However, reward hacking, where agents maximize a s…