6 papers
S-EMBER: A Large-Scale Benchmark for Streaming Egocentric Memory Retrieval
Xiaodong Wang, Xuanyi Zhao, Pedro Rodriguez +7
As wearable devices enable continuous first-person recording, AI assistants must reason across long time horizons to recall past experiences-a capability known as episodic memory.…
HoneyBee: Data Recipes for Vision-Language Reasoners
Hritik Bansal, Devendra Singh Sachan, Kai-Wei Chang +4
Recent advances in vision-language models (VLMs) have made them highly effective at reasoning tasks. However, the principles underlying the construction of performant VL reasoning…
Continual Learning via Sparse Memory Finetuning
Jessy Lin, Luke Zettlemoyer, Gargi Ghosh +4
Modern language models are powerful, but typically static after deployment. A major obstacle to building models that continually learn over time is catastrophic forgetting, where u…
Learning Facts at Scale with Active Reading
Jessy Lin, Vincent-Pierre Berges, Xilun Chen +3
LLMs are known to store vast amounts of knowledge in their parametric memory. However, learning and recalling facts from this memory is known to be unreliable, depending largely on…
Learning to Reason for Factuality
Xilun Chen, Ilia Kulikov, Vincent-Pierre Berges +5
Reasoning Large Language Models (R-LLMs) have significantly advanced complex reasoning tasks but often struggle with factuality, generating substantially more hallucinations than t…
FACTORY: A Challenging Human-Verified Prompt Set for Long-Form Factuality
Mingda Chen, Yang Li, Xilun Chen +3
Long-form factuality evaluation assesses the ability of models to generate accurate, comprehensive responses to short prompts. Existing benchmarks often lack human verification, le…