1 citations · 1 across the 2 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
ASCIIEval: Benchmarking Models' Visual Perception in Text Strings via ASCII Art
Qi Jia, Xiang Yue, Shanshan Huang +5
Perceiving visual semantics embedded within consecutive characters is a crucial yet under-explored capability for both Large Language Models (LLMs) and Multi-modal Large Language M…
cs.CL2024
SimulBench: Evaluating Language Models with Creative Simulation Tasks
Qi Jia, Xiang Yue, Tianyu Zheng +2
We introduce SimulBench, a benchmark designed to evaluate large language models (LLMs) across a diverse collection of creative simulation scenarios, such as acting as a Linux termi…