41 citations · 89 across the 31 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Probing Memes in LLMs: A Paradigm for the Entangled Evaluation World
Luzhou Peng, Zhengxin Yang, Honglu Ji +6
Current evaluation paradigms for large language models (LLMs) characterize models and datasets separately, yielding coarse descriptions: items in datasets are treated as pre-labele…
cs.CL2023
AGIBench: A Multi-granularity, Multimodal, Human-referenced, Auto-scoring Benchmark for Large Language Models
Fei Tang, Wanling Gao, Luzhou Peng +1
Large language models (LLMs) like ChatGPT have revealed amazing intelligence. How to evaluate the question-solving abilities of LLMs and their degrees of intelligence is a hot-spot…