Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
SOP-Maze: Evaluating Large Language Models on Complicated Business Standard Operating Procedures
Jiaming Wang, Zhe Tang, Zehao Jin +5
As large language models (LLMs) are widely deployed as domain-specific agents, many benchmarks have been proposed to evaluate their ability to follow instructions and make decision…
cs.CL2025
Meeseeks: A Feedback-Driven, Iterative Self-Correction Benchmark evaluating LLMs' Instruction Following Capability
Jiaming wang, Yunke Zhao, Peng Ding +8
The capability to precisely adhere to instructions is a cornerstone for Large Language Models (LLMs) to function as dependable agents in real-world scenarios. However, confronted w…
cs.CL2025
SDGO: Self-Discrimination-Guided Optimization for Consistent Safety in Large Language Models
Peng Ding, Wen Sun, Dailin Li +4
Large Language Models (LLMs) excel at various natural language processing tasks but remain vulnerable to jailbreaking attacks that induce harmful content generation. In this paper,…