1 paper
Christodoulos Constantinides, Dhaval Patel, Shuxin Lin +3
We introduce FailureSensorIQ, a novel Multi-Choice Question-Answering (MCQA) benchmarking system designed to assess the ability of Large Language Models (LLMs) to reason and unders…