activity
20242026
collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2026

Experiments or Outcomes? Probing Scientific Feasibility in Large Language Models

Seyedali Mohammadi, Manas Gaur, Francis Ferraro

Scientific feasibility assessment asks whether a claim is consistent with established knowledge and whether experimental evidence could support or refute it. We frame feasibility a…

cs.CL2026

Flying Pigs, FaR and Beyond: Evaluating LLM Reasoning in Counterfactual Worlds

Anish R Joishy, Ishwar B Balappanawar, Vamshi Krishna Bonagiri +3

A fundamental challenge in reasoning is navigating hypothetical, counterfactual worlds where logic may conflict with ingrained knowledge. We investigate this frontier for Large Lan…

cs.CL2026

Beyond Memorization: Testing LLM Reasoning on Unseen Theory of Computation Tasks

Shlok Shelat, Jay Raval, Souvik Roy +1

Large language models (LLMs) have demonstrated strong performance on formal language tasks, yet whether this reflects genuine symbolic reasoning or pattern matching on familiar con…

cs.CL2025

SymLoc: Symbolic Localization of Hallucination across HaluEval and TruthfulQA

Naveen Lamba, Sanju Tiwari, Manas Gaur

LLMs still struggle with hallucination, especially when confronted with symbolic triggers like modifiers, negation, numbers, exceptions, and named entities. Yet, we lack a clear un…

cs.CL2025

Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA

Naveen Lamba, Sanju Tiwari, Manas Gaur

Hallucination in Large Language Models (LLMs) is a well studied problem. However, the properties that make LLM intrinsically vulnerable to hallucinations have not been identified a…

cs.CL2025

Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives

Batool Haider, Atmika Gorti, Aman Chadha +1

Large Language Models (LLMs) in mental healthcare risk propagating biases that reinforce stigma and harm marginalized groups. While previous research identified concerning trends,…