3 papers
cs.AI2026
StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models
Jinghan Tan, Yuanzheng Wang, Lu Chen +3
As large language models are increasingly used in data-scarce and evolving task scenarios, few-shot in-context learning (ICL) has become a key paradigm for task adaptation. However…
cs.CL2026
DocAtlas: Long-Document Understanding as Mutable-State Interaction
Hongchen Wei, Yuanzhe Wang, Bei Liu +8
Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts. Existing retrieval-augmented systems usually selec…
cs.CL2026
XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding
Hongchen Wei, Yuanzhe Wang, Bei Liu +9
Real-world document tasks often ask professionals to answer questions from annual reports, regulations, clinical guidelines, and technical manuals that span hundreds or thousands o…