1 paper · 1 filter
Belinda Z. Li, Been Kim, Zi Wang
Large language models (LLMs) have shown impressive performance on reasoning benchmarks like math and logic. While many works have largely assumed well-defined tasks, real-world que…