3 papers
cs.AI2026
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy
Kazem Faghih, Yize Cheng, Shoumik Saha +3
Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in differe…
cs.CV2025
SliderEdit: Continuous Image Editing with Fine-Grained Instruction Control
Arman Zarei, Samyadeep Basu, Mobina Pournemat +3
Instruction-based image editing models have recently achieved impressive performance, enabling complex edits to an input image from a multi-instruction prompt. However, these model…
cs.CL2025
Reasoning Under Uncertainty: Exploring Probabilistic Reasoning Capabilities of LLMs
Mobina Pournemat, Keivan Rezaei, Gaurang Sriramanan +5
Despite widespread success in language understanding and generation, large language models (LLMs) exhibit unclear and often inconsistent behavior when faced with tasks that require…