11 papers
TS-Haystack: A Multi-Task Retrieval Benchmark for Long-Context Time-Series Reasoning
Nicolas Zumarraga, Thomas Kaar, Ning Wang +12
Time Series Language Models (TSLMs) promise reasoning over real-world temporal data, but their ability to retrieve and reason over long time-series remains largely untested. We int…
LatentGym: A Testbed For Cross-Task Experiential Learning With Controllable Latent Structure
Daksh Mittal, Tommaso Castellani, Thomson Yen +7
We envision continually learning agentic systems that become more useful over time: as they encounter sequences of related tasks, they should infer the hidden structure shared acro…
Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards
Fang Wu, Aaron Tu, Weihao Xuan +21
Reinforcement learning with verifiable rewards (RLVR) is a practical, scalable way to improve large language models on math, code, and other structured tasks. However, we argue tha…
ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents
Yuxiang Lai, Peng Xia, Haonian Ji +8
Interactive agent benchmarks face a tension between scalable construction and realistic workflow evaluation. Hand-authored tasks are expensive to extend and revise, while static pr…
BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning
Yuhang Xu, Kaibin Tian, Yang Tian +6
Reinforcement Learning (RL) has become a cornerstone for improving the performance of Large Language Models (LLMs). However, its rollout phase constitutes a significant efficiency…
NoiseRater: Meta-Learned Noise Valuation for Diffusion Model Training
Fang Wu, Haokai Zhao, Da Xing +17
Diffusion models have achieved remarkable success across a wide range of generative tasks, yet their training paradigm largely treats injected noise as uniformly informative. In th…