7 citations · 37 across the 24 of their papers we have counts for
44 papers
Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
Jacqueline He, Howard Yen, Shuyue Stella Li +9
Logit-based knowledge distillation (KD) is used to train smaller language models (LMs) via supervision from stronger teachers, but whether its benefits are consistent across traini…
S-EMBER: A Large-Scale Benchmark for Streaming Egocentric Memory Retrieval
Xiaodong Wang, Xuanyi Zhao, Pedro Rodriguez +7
As wearable devices enable continuous first-person recording, AI assistants must reason across long time horizons to recall past experiences-a capability known as episodic memory.…
Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility
Gergely Szilvasy, Manuel Faysse, Maria Lomeli +5
Under modern test-time compute and agentic paradigms, language models process ever-longer sequences. Efficient text generation with transformer architectures is increasingly constr…
Procedural Knowledge at Scale Improves Reasoning
Di Wu, Devendra Singh Sachan, Wen-tau Yih +1
Test-time scaling has emerged as an effective way to improve language models on challenging reasoning tasks. However, most existing methods treat each problem in isolation and do n…
Anchored Decoding: Provably Reducing Copyright Risk for Any Language Model
Jacqueline He, Jonathan Hayase, Wen-tau Yih +3
Language models (LMs) tend to memorize portions of their training data and emit verbatim spans. When the underlying sources are sensitive or copyright-protected, such reproduction…
DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
Rulin Shao, Akari Asai, Shannon Zejiang Shen +18
Deep research agents perform multi-step research to produce long-form, well-attributed answers. However, most open deep research agents are trained on easily verifiable short-form…