19 papers
MemFail: Stress-Testing Failure Modes of LLM Memory Systems
Ishir Garg, Neel Kolhe, Dawn Song +1
Large language model (LLM) agents increasingly rely on external memory systems to remain consistent across long-horizon interactions, but little empirical work has been done to und…
InfoSynth: Information-Guided Benchmark Synthesis for LLMs
Ishir Garg, Neel Kolhe, Xuandong Zhao +1
Large language models (LLMs) have demonstrated significant advancements in reasoning and code generation, but efficiently creating new benchmarks to evaluate these capabilities rem…
Learning to Reason without External Rewards
Xuandong Zhao, Zhewei Kang, Aosong Feng +2
Training large language models (LLMs) for complex reasoning via Reinforcement Learning with Verifiable Rewards (RLVR) is effective but limited by reliance on costly, domain-specifi…
Position: LLM Watermarking Should Align Stakeholders' Incentives for Practical Adoption
Yepeng Liu, Xuandong Zhao, Dawn Song +2
Despite progress in watermarking algorithms for large language models (LLMs), real-world deployment remains limited. We argue that this gap stems from misaligned incentives among L…
In-Context Watermarks for Large Language Models
Yepeng Liu, Xuandong Zhao, Christopher Kruegel +2
The growing use of large language models (LLMs) for sensitive applications has highlighted the need for effective watermarking techniques to ensure the provenance and accountabilit…
AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents
Jingxu Xie, Dylan Xu, Xuandong Zhao +1
We introduce AgentSynth, a scalable and cost-efficient pipeline for automatically synthesizing high-quality tasks and trajectory datasets for generalist computer-use agents. Levera…