3 papers
cs.LG2026
Beyond Naïve Prompting: Strategies for Improved Context-aided Forecasting with LLMs
Arjun Ashok, Andrew Robert Williams, Vincent Zhihao Zheng +5
Real-world forecasting requires models to integrate not only historical data but also relevant contextual information provided in textual form. While large language models (LLMs) s…
cs.CL2026
DRBench: A Realistic Benchmark for Enterprise Deep Research
Amirhossein Abaskohi, Tianyi Chen, Miguel Muñoz-Mármol +11
We introduce DRBench, a benchmark for evaluating AI agents on complex, open-ended deep research tasks in enterprise settings. Unlike prior benchmarks that focus on simple questions…
cs.LG2025
Context is Key: A Benchmark for Forecasting with Essential Textual Information
Andrew Robert Williams, Arjun Ashok, Ãtienne Marcotte +8
Forecasting is a critical task in decision-making across numerous domains. While historical numerical data provide a start, they fail to convey the complete context for reliable an…