4 papers · 1 filter
ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints
Hyeonje Choi, Jeongsoo Lee, Hyojun Lee +1
We introduce \ToolMATH, a math-grounded diagnostic benchmark for evaluating long-horizon tool use under controllable tool-catalog conditions. \ToolMATH converts stepwise MATH solut…
Latent Self-Consistency for Reliable Majority-Set Selection in Short- and Long-Answer Reasoning
Jungsuk Oh, Jay-Yoon Lee
Probabilistic decoding in Large Language Models (LLMs) often yields inconsistent outputs, particularly on complex or long-form questions. Self-Consistency (SC) mitigates this for s…
LaSEr-Edit: Localized Span-level Error Editing with Energy-based Localization
Hye Ryung Son, Jay-Yoon Lee, Saehee Eom +1
As large language models (LLMs) are widely adopted in real-world applications, it has become critical to ensure LLMs satisfy safety constraints, such as non-toxicity and logical co…
Case-Based Reasoning Approach for Solving Financial Question Answering
Yikyung Kim, Jay-Yoon Lee
Measuring a machine's understanding of human language often involves assessing its reasoning skills, i.e. logical process of deriving answers to questions. While recent language mo…