Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
LongIns: A Challenging Long-context Instruction-based Exam for LLMs
Shawn Gavin, Tuney Zheng, Jiaheng Liu +6
The long-context capabilities of large language models (LLMs) have been a hot topic in recent years. To evaluate the performance of LLMs in different scenarios, various assessment…
cs.CL2024
ING-VP: MLLMs cannot Play Easy Vision-based Games Yet
Haoran Zhang, Hangyu Guo, Shuyue Guo +4
As multimodal large language models (MLLMs) continue to demonstrate increasingly competitive performance across a broad spectrum of tasks, more intricate and comprehensive benchmar…
cs.CL2024
HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models
Haoran Que, Feiyu Duan, Liqun He +11
In recent years, Large Language Models (LLMs) have demonstrated remarkable capabilities in various tasks (e.g., long-context understanding), and many benchmarks have been proposed.…