1 citations · 1 across the 2 of their papers we have counts for
4 papers
SOP-Maze: Evaluating Large Language Models on Complicated Business Standard Operating Procedures
Jiaming Wang, Zhe Tang, Zehao Jin +5
As large language models (LLMs) are widely deployed as domain-specific agents, many benchmarks have been proposed to evaluate their ability to follow instructions and make decision…
SDGO: Self-Discrimination-Guided Optimization for Consistent Safety in Large Language Models
Peng Ding, Wen Sun, Dailin Li +4
Large Language Models (LLMs) excel at various natural language processing tasks but remain vulnerable to jailbreaking attacks that induce harmful content generation. In this paper,…
Vaiage: A Multi-Agent Solution to Personalized Travel Planning
Binwen Liu, Jiexi Ge, Jiamin Wang
Planning trips is a cognitively intensive task involving conflicting user preferences, dynamic external information, and multi-step temporal-spatial optimization. Traditional platf…
Meeseeks: A Feedback-Driven, Iterative Self-Correction Benchmark evaluating LLMs' Instruction Following Capability
Jiaming wang, Yunke Zhao, Peng Ding +8
The capability to precisely adhere to instructions is a cornerstone for Large Language Models (LLMs) to function as dependable agents in real-world scenarios. However, confronted w…