2 papers
cs.SE2025
Rethinking Testing for LLM Applications: Characteristics, Challenges, and a Lightweight Interaction Protocol
Wei Ma, Yixiao Yang, Qiang Hu +8
Applications of Large Language Models~(LLMs) have evolved from simple text generators into complex software systems that integrate retrieval augmentation, tool invocation, and mult…
cs.AI2025
Importance Sampling is All You Need: Predict LLM's performance on new benchmark by reusing existing benchmark
Junjie Shi, Wei Ma, Shi Ying +3
With the rapid advancement of large language models , code generation has become a key benchmark for evaluating LLM capabilities. However, existing benchmarks face two major challe…