2 papers
cs.AI2026
Redundant or Necessary? A Benchmark for Detecting Redundant Steps in Agent Trajectories
Minyang Hu, Bo Yang, Zhinuo Zhou +4
LLM-based agents have demonstrated strong capabilities in solving complex tasks through multi-step reasoning and tool use. However, existing evaluation protocols primarily focus on…
cs.AI2026
Opt-Verifier: Unleashing the Power of LLMs for Optimization Modeling via Dual-Side Verification
Haoyang Liu, Jie Wang, Boxuan Niu +8
Building mathematical optimization models is critical in operations research (OR), while it requires substantial human expertise. Recent advancements have utilized large language m…