5 papers
ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints
Hyeonje Choi, Jeongsoo Lee, Hyojun Lee +1
We introduce \ToolMATH, a math-grounded diagnostic benchmark for evaluating long-horizon tool use under controllable tool-catalog conditions. \ToolMATH converts stepwise MATH solut…
Latent Self-Consistency for Reliable Majority-Set Selection in Short- and Long-Answer Reasoning
Jungsuk Oh, Jay-Yoon Lee
Probabilistic decoding in Large Language Models (LLMs) often yields inconsistent outputs, particularly on complex or long-form questions. Self-Consistency (SC) mitigates this for s…
UniFault: A Fault Diagnosis Foundation Model from Bearing Data
Emadeldeen Eldele, Mohamed Ragab, Xu Qing +5
Machine fault diagnosis (FD) is a critical task for predictive maintenance, enabling early fault detection and preventing unexpected failures. Despite its importance, existing FD m…
Stop-RAG: Value-Based Retrieval Control for Iterative RAG
Jaewan Park, Solbee Cho, Jay-Yoon Lee
Iterative retrieval-augmented generation (RAG) enables large language models to answer complex multi-hop questions, but each additional loop increases latency, costs, and the risk…
Retargeting an Abstract Interpreter for a New Language by Partial Evaluation
Jay Lee
It is well-known that abstract interpreters can be systematically derived from their concrete counterparts using a "recipe," but developing sound static analyzers remains a time-co…