1 paper · 1 filter
Abhishek Chandwani, Ishan Gupta
Large language models excel on objectively verifiable tasks such as math and programming, where evaluation reduces to unit tests or a single correct answer. In contrast, real-world…