1 paper · 1 filter
Hongbo Jin, Chi Wang, Haoran Tang +5
Recent benchmarks reveal that despite strong reasoning capabilities, large language models (LLMs) still struggle to faithfully apply complex contextual knowledge. These failures ar…