Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
RepBench: Compiling Benchmarks into Capability Representations for Large Language Models
Yanshi Li, Xueru Bai, Shuman Liu +1
Representation engineering reads and steers capability directions in large language models, yet methods are typically evaluated on paper-specific synthetic data. The resulting meas…
cs.CL2025
WideSearch: Benchmarking Agentic Broad Info-Seeking
Ryan Wong, Jiawei Wang, Junjie Zhao +10
From professional research to everyday planning, many tasks are bottlenecked by wide-scale information seeking, which is more repetitive than cognitively complex. With the rapid de…
cs.CL2025
LFD: Layer Fused Decoding to Exploit External Knowledge in Retrieval-Augmented Generation
Yang Sun, Zhiyong Xie, Lixin Zou +7
Retrieval-augmented generation (RAG) incorporates external knowledge into large language models (LLMs), improving their adaptability to downstream tasks and enabling information up…