collaborators

11 papers

cs.LG2026

TS-Haystack: A Multi-Task Retrieval Benchmark for Long-Context Time-Series Reasoning

Nicolas Zumarraga, Thomas Kaar, Ning Wang +12

Time Series Language Models (TSLMs) promise reasoning over real-world temporal data, but their ability to retrieve and reason over long time-series remains largely untested. We int…

cs.LG2026

LatentGym: A Testbed For Cross-Task Experiential Learning With Controllable Latent Structure

Daksh Mittal, Tommaso Castellani, Thomson Yen +7

We envision continually learning agentic systems that become more useful over time: as they encounter sequences of related tasks, they should infer the hidden structure shared acro…

cs.LG2026

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards

Fang Wu, Aaron Tu, Weihao Xuan +21

Reinforcement learning with verifiable rewards (RLVR) is a practical, scalable way to improve large language models on math, code, and other structured tasks. However, we argue tha…

cs.AI2026

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents

Yuxiang Lai, Peng Xia, Haonian Ji +8

Interactive agent benchmarks face a tension between scalable construction and realistic workflow evaluation. Hand-authored tasks are expensive to extend and revise, while static pr…

cs.LG2026

BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning

Yuhang Xu, Kaibin Tian, Yang Tian +6

Reinforcement Learning (RL) has become a cornerstone for improving the performance of Large Language Models (LLMs). However, its rollout phase constitutes a significant efficiency…

cs.LG2026

NoiseRater: Meta-Learned Noise Valuation for Diffusion Model Training

Fang Wu, Haokai Zhao, Da Xing +17

Diffusion models have achieved remarkable success across a wide range of generative tasks, yet their training paradigm largely treats injected noise as uniformly informative. In th…