1 paper · 1 filter
Ali Khoramfar, Ali Ramezani, Mohammad Mahdi Mohajeri +3
While Large Language Models (LLMs) achieve near-human performance on standard benchmarks, their capabilities often fail to generalize to complex, real-world problems. To bridge thi…