3 papers
cs.LG2025
ChessQA: Evaluating Large Language Models for Chess Understanding
Qianfeng Wen, Zhenwei Tang, Ashton Anderson
Chess provides an ideal testbed for evaluating the reasoning, modeling, and abstraction capabilities of large language models (LLMs), as it has well-defined structure and objective…
cs.AI2025
SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
Zhenwei Tang, Difan Jiao, Blair Yang +1
Evaluating whether vision-language models (VLMs) reason consistently across representations is challenging because modality comparisons are typically confounded by task differences…
cs.AI2025
Learning to Imitate with Less: Efficient Individual Behavior Modeling in Chess
Zhenwei Tang, Difan Jiao, Eric Xue +4
As humans seek to collaborate with, learn from, and better understand artificial intelligence systems, developing AIs that can accurately emulate individual decision-making becomes…