29 citations · 86 across the 15 of their papers we have counts for
31 papers · 1 filter
MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding
Bosi Wen, Cunxiang Wang, Jiayi Gui +6
Recently, the rapid development of large language models (LLMs) has reshaped software engineering by enabling autonomous code agents that plan, execute, and utilize external tools…
IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation
Bosi Wen, Yilin Niu, Cunxiang Wang +5
Instruction-following is a foundational capability of large language models (LLMs), with its improvement hinging on scalable and accurate feedback from judge models. However, the r…
IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation
Bosi Wen, Yilin Niu, Cunxiang Wang +6
Instruction-following is a fundamental ability of Large Language Models (LLMs), requiring their generated outputs to follow multiple constraints imposed in input instructions. Nume…
Training Language Model to Critique for Better Refinement
Tianshu Yu, Chao Xiang, Mingchuan Yang +8
Large language models (LLMs) have demonstrated remarkable evaluation and critique capabilities, providing insightful feedback and identifying flaws in various tasks. However, limit…
HPSS: Heuristic Prompting Strategy Search for LLM Evaluators
Bosi Wen, Pei Ke, Yufei Sun +6
Since the adoption of large language models (LLMs) for text evaluation has become increasingly prevalent in the field of natural language processing (NLP), a series of existing wor…
The Superalignment of Superhuman Intelligence with Large Language Models
Minlie Huang, Yingkang Wang, Shiyao Cui +2
We have witnessed superhuman intelligence thanks to the fast development of large language models and multimodal language models. As the application of such superhuman models becom…