collaborators

10 papers

cs.CL2026

Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search

Howard Yen, Yoonsang Lee, Ashwin Paranjape +5

The paper introduces SLIM, a lightweight framework that separates search and browsing tools and periodically summarizes information to overcome context limits in long-horizon web‑a…

cs.CL2026

DySCO: Dynamic Attention-Scaling Decoding for Long-Context Language Models

Xi Ye, Wuwei Zhang, Fangcong Yin +2

Understanding and reasoning over long contexts is a crucial capability for language models (LMs). Although recent models support increasingly long context windows, their accuracy o…

cs.CL2026

Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasks

Yoonsang Lee, Howard Yen, Xi Ye +1

We study parallel test-time scaling for long-horizon agentic tasks such as agentic search and deep research, where multiple rollouts are generated in parallel and aggregated into a…

cs.CL2025

How to Train Long-Context Language Models (Effectively)

Tianyu Gao, Alexander Wettig, Howard Yen +1

We study continued training and supervised fine-tuning (SFT) of a language model (LM) to make effective use of long-context information. We first establish a reliable evaluation pr…

cs.CL2025

MMTEB: Massive Multilingual Text Embedding Benchmark

Kenneth Enevoldsen, Isaac Chung, Imene Kerboua +83

Text embeddings are typically evaluated on a limited set of tasks, which are constrained by language, domain, and task diversity. To address these limitations and provide a more co…

cs.CL2025

Precise Information Control in Long-Form Text Generation

Jacqueline He, Howard Yen, Margaret Li +7

A central challenge in language models (LMs) is faithfulness hallucination: the generation of information unsubstantiated by input context. To study this problem, we propose Precis…