Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
ConInstruct: Evaluating Large Language Models on Conflict Detection and Resolution in Instructions
Xingwei He, Qianru Zhang, Pengfei Chen +4
Instruction-following is a critical capability of Large Language Models (LLMs). While existing works primarily focus on assessing how well LLMs adhere to user instructions, they of…
cs.CL2025
EffiBench-X: A Multi-Language Benchmark for Measuring Efficiency of LLM-Generated Code
Yuhao Qing, Boyu Zhu, Mingzhe Du +9
Existing code generation benchmarks primarily evaluate functional correctness, with limited focus on code efficiency and often restricted to a single language like Python. To addre…