benchmark compilation 1capability probing 1evaluation methods 1large language models 1representation learning 1
From the 1 of 7 linked papers with an AI index.
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
RepBench: Compiling Benchmarks into Capability Representations for Large Language Models
Yanshi Li, Xueru Bai, Shuman Liu +1
The paper introduces RepBench, a framework that aggregates thousands of benchmark datasets into a large set of probe texts to evaluate capability-aligned representations of large l…
cs.CL2026
LFD: Layer Fused Decoding to Exploit External Knowledge in Retrieval-Augmented Generation
Yang Sun, Zhiyong Xie, Lixin Zou +7
Retrieval-augmented generation (RAG) incorporates external knowledge into large language models (LLMs), improving their adaptability to downstream tasks and enabling information up…
cs.CL2025
WideSearch: Benchmarking Agentic Broad Info-Seeking
Ryan Wong, Jiawei Wang, Junjie Zhao +10
From professional research to everyday planning, many tasks are bottlenecked by wide-scale information seeking, which is more repetitive than cognitively complex. With the rapid de…