1 paper · 1 filter
Andrew Blair-Stanek, Benjamin Van Durme
An LLM is stable if it reaches the same conclusion when asked the identical question multiple times. We find leading LLMs like gpt-4o, claude-3.5, and gemini-1.5 are unstable when…