collaborators

22 papers

cs.AI2026

SciR: A Controllable Benchmark for Scientific Reasoning in LLMs

Pierre Beckmann, Marco Valentino, Andre Freitas

Three paradigmatic forms of inference recur across scientific reasoning: deduction, induction, and causal abduction. Reliably evaluating LLMs on these in scientific settings is cur…

cs.AI2026

Neurosymbolic Clinical Trial Matching via LLM-Driven Abduction and Logical Verification

Baiyang Qu, Leonardo Ranaldi, Xi Wang +1

Large Language Models (LLMs) offer a promising path to automate Clinical Trial Matching (CTM), but still struggle with the deterministic verification required for complex eligibili…

cs.CL2026

Compliance versus Sensibility: On the Reasoning Controllability in Large Language Models

Xingwei Tan, Marco Valentino, Mahmud Elahi Akhter +3

Large Language Models (LLMs) are known to acquire reasoning capabilities through shared inference patterns in pre-training data, which are further elicited via Chain-of-Thought (Co…

cs.CL2026

Is Inference Mediated by Distinct Semantic Structures in LLMs? A Mechanistic Interpretation

Nura Aljaafari, Marco Valentino, André Freitas

Predicting a label correctly does not necessarily require representing the operation that produces it. Transformer representations are known to carry label-level information, but w…

cs.CL2026

Monotonic Reference-Free Refinement for Autoformalization

Lan Zhang, Marco Valentino, André Freitas

While statement autoformalization has advanced rapidly, full-theorem autoformalization remains largely unexplored. Existing iterative refinement methods in statement autoformalizat…

cs.AI2026

Mitigating Content Effects on Reasoning in Language Models through Fine-Grained Activation Steering

Marco Valentino, Geonhee Kim, Dhairya Dalal +2

Large language models (LLMs) exhibit reasoning biases, often conflating content plausibility with formal logical validity. This can lead to wrong inferences in critical domains, wh…