Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
HammerBench: Fine-Grained Function-Calling Evaluation in Real Mobile Device Scenarios
Jun Wang, Jiamu Zhou, Muning Wen +7
Evaluating the performance of LLMs in multi-turn human-agent interactions presents significant challenges, particularly due to the complexity and variability of user behavior. In t…
cs.CL2024
P3: A Policy-Driven, Pace-Adaptive, and Diversity-Promoted Framework for data pruning in LLM Training
Yingxuan Yang, Huayi Wang, Muning Wen +4
In the rapidly advancing field of Large Language Models (LLMs), effectively leveraging existing datasets during fine-tuning to maximize the model's potential is of paramount import…