From the 1 of 13 linked papers with an AI index.
13 papers
Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search
Howard Yen, Yoonsang Lee, Ashwin Paranjape +5
The paper introduces SLIM, a lightweight framework that separates search and browsing tools and periodically summarizes information to overcome context limits in long-horizon web‑a…
Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision
Yinghui He, Simran Kaur, Adithya Bhaskar +7
Current post-training methods in verifiable settings fall into two categories. Reinforcement learning (RLVR) relies on binary rewards, which are broadly applicable and powerful, bu…
Hubs or Fringes: Pretraining Data Selection via Web Graph Centrality
Vedant Badoni, Danqi Chen, Xinyi Wang
The performance of modern language models depends critically on pretraining data composition. Yet existing data selection methods rely on auxiliary classifiers for document scoring…
DySCO: Dynamic Attention-Scaling Decoding for Long-Context Language Models
Xi Ye, Wuwei Zhang, Fangcong Yin +2
Understanding and reasoning over long contexts is a crucial capability for language models (LMs). Although recent models support increasingly long context windows, their accuracy o…
Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasks
Yoonsang Lee, Howard Yen, Xi Ye +1
We study parallel test-time scaling for long-horizon agentic tasks such as agentic search and deep research, where multiple rollouts are generated in parallel and aggregated into a…
How to Train Long-Context Language Models (Effectively)
Tianyu Gao, Alexander Wettig, Howard Yen +1
We study continued training and supervised fine-tuning (SFT) of a language model (LM) to make effective use of long-context information. We first establish a reliable evaluation pr…