8 papers · 1 filter
Experiments or Outcomes? Probing Scientific Feasibility in Large Language Models
Seyedali Mohammadi, Manas Gaur, Francis Ferraro
Scientific feasibility assessment asks whether a claim is consistent with established knowledge and whether experimental evidence could support or refute it. We frame feasibility a…
Flying Pigs, FaR and Beyond: Evaluating LLM Reasoning in Counterfactual Worlds
Anish R Joishy, Ishwar B Balappanawar, Vamshi Krishna Bonagiri +3
A fundamental challenge in reasoning is navigating hypothetical, counterfactual worlds where logic may conflict with ingrained knowledge. We investigate this frontier for Large Lan…
Beyond Memorization: Testing LLM Reasoning on Unseen Theory of Computation Tasks
Shlok Shelat, Jay Raval, Souvik Roy +1
Large language models (LLMs) have demonstrated strong performance on formal language tasks, yet whether this reflects genuine symbolic reasoning or pattern matching on familiar con…
SymLoc: Symbolic Localization of Hallucination across HaluEval and TruthfulQA
Naveen Lamba, Sanju Tiwari, Manas Gaur
LLMs still struggle with hallucination, especially when confronted with symbolic triggers like modifiers, negation, numbers, exceptions, and named entities. Yet, we lack a clear un…
Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA
Naveen Lamba, Sanju Tiwari, Manas Gaur
Hallucination in Large Language Models (LLMs) is a well studied problem. However, the properties that make LLM intrinsically vulnerable to hallucinations have not been identified a…
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
Batool Haider, Atmika Gorti, Aman Chadha +1
Large Language Models (LLMs) in mental healthcare risk propagating biases that reinforce stigma and harm marginalized groups. While previous research identified concerning trends,…