22 papers
SciR: A Controllable Benchmark for Scientific Reasoning in LLMs
Pierre Beckmann, Marco Valentino, Andre Freitas
Three paradigmatic forms of inference recur across scientific reasoning: deduction, induction, and causal abduction. Reliably evaluating LLMs on these in scientific settings is cur…
Neurosymbolic Clinical Trial Matching via LLM-Driven Abduction and Logical Verification
Baiyang Qu, Leonardo Ranaldi, Xi Wang +1
Large Language Models (LLMs) offer a promising path to automate Clinical Trial Matching (CTM), but still struggle with the deterministic verification required for complex eligibili…
Compliance versus Sensibility: On the Reasoning Controllability in Large Language Models
Xingwei Tan, Marco Valentino, Mahmud Elahi Akhter +3
Large Language Models (LLMs) are known to acquire reasoning capabilities through shared inference patterns in pre-training data, which are further elicited via Chain-of-Thought (Co…
Is Inference Mediated by Distinct Semantic Structures in LLMs? A Mechanistic Interpretation
Nura Aljaafari, Marco Valentino, André Freitas
Predicting a label correctly does not necessarily require representing the operation that produces it. Transformer representations are known to carry label-level information, but w…
Monotonic Reference-Free Refinement for Autoformalization
Lan Zhang, Marco Valentino, André Freitas
While statement autoformalization has advanced rapidly, full-theorem autoformalization remains largely unexplored. Existing iterative refinement methods in statement autoformalizat…
Mitigating Content Effects on Reasoning in Language Models through Fine-Grained Activation Steering
Marco Valentino, Geonhee Kim, Dhairya Dalal +2
Large language models (LLMs) exhibit reasoning biases, often conflating content plausibility with formal logical validity. This can lead to wrong inferences in critical domains, wh…