10 papers
GTS: Inference-Time Scaling of Latent Reasoning with a Learnable Gaussian Thought Sampler
Minghan Wang, Ye Bai, Thuy-Trang Vu +2
Inference-time scaling (ITS) in latent reasoning models typically relies on heuristic perturbations, such as dropout or fixed Gaussian noise, to generate diverse candidate trajecto…
Improving Symbolic Translation of Language Models for Logical Reasoning
Ramya Keerthy Thatikonda, Jiuzhou Han, Wray Buntine +1
The use of formal language for deductive logical reasoning aligns well with language models (LMs), where translating natural language (NL) into first-order logic (FOL) and employin…
Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
Dhruv Anand, Ehsan Shareghi
We introduce Cube Bench, a Rubik's-cube benchmark for evaluating spatial and sequential reasoning in multimodal large language models (MLLMs). The benchmark decomposes performance…
Towards Inference-time Scaling for Continuous Space Reasoning
Minghan Wang, Thuy-Trang Vu, Ehsan Shareghi +1
Inference-time scaling through multiple sample generation in combination with Process- or Outcome-Reward Model (PRM or ORM) re-ranking has proven effective for text-based reasoning…
Assessing the Sensitivity and Alignment of FOL Closeness Metrics
Ramya Keerthy Thatikonda, Wray Buntine, Ehsan Shareghi
The recent successful paradigm of solving logical reasoning problems with tool-augmented large language models (LLMs) leverages translation of natural language (NL) statements into…
Logical Reasoning with Outcome Reward Models for Test-Time Scaling
Ramya Keerthy Thatikonda, Wray Buntine, Ehsan Shareghi
Logical reasoning is a critical benchmark for evaluating the capabilities of large language models (LLMs), as it reflects their ability to derive valid conclusions from given premi…