activity
20242026
collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2025

CFBench: A Comprehensive Constraints-Following Benchmark for LLMs

Tao Zhang, Chenglin Zhu, Yanjun Shen +10

The adeptness of Large Language Models (LLMs) in comprehending and following natural language instructions is critical for their deployment in sophisticated real-world applications…

cs.CL2025

Facilitating Multi-turn Function Calling for LLMs via Compositional Instruction Tuning

Mingyang Chen, Haoze Sun, Tianpeng Li +7

Large Language Models (LLMs) have exhibited significant potential in performing diverse tasks, including the ability to call functions or use external tools to enhance their perfor…

cs.CL2025

MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification

Linzhuang Sun, Hao Liang, Jingxuan Wei +5

According to the Test-Time Scaling, the integration of External Slow-Thinking with the Verify mechanism has been demonstrated to enhance multi-round reasoning in large language mod…

cs.CL2025

FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human Feedback

Youquan Li, Miao Zheng, Fan Yang +5

Human feedback is crucial in the interactions between humans and Large Language Models (LLMs). However, existing research primarily focuses on benchmarking LLMs in single-turn dial…

cs.CL2024

BaichuanSEED: Sharing the Potential of ExtensivE Data Collection and Deduplication by Introducing a Competitive Large Language Model Baseline

Guosheng Dong, Da Pan, Yiding Sun +17

The general capabilities of Large Language Models (LLM) highly rely on the composition and selection on extensive pretraining datasets, treated as commercial secrets by several ins…

cs.CL2024

PAS: Data-Efficient Plug-and-Play Prompt Augmentation System

Miao Zheng, Hao Liang, Fan Yang +16

In recent years, the rise of Large Language Models (LLMs) has spurred a growing demand for plug-and-play AI systems. Among the various AI techniques, prompt engineering stands out…