2 papers
cs.AI2026
AgentPersonaBench: Benchmarking Persona-Driven User Simulation
Jintao Huang, Yifan Wang, Hongyu Shen +43
We introduce AgentPersonaBench (APB), a benchmark evaluating whether persona conditioning faithfully steers downstream agent behavior. While language models are increasingly deploy…
cs.CV2026
KazakhOCR: A Synthetic Benchmark for Evaluating Multimodal Models in Low-Resource Kazakh Script OCR
Henry Gagnier, Sophie Gagnier, Ashwin Kirubakaran
Kazakh is a Turkic language using the Arabic, Cyrillic, and Latin scripts, making it unique in terms of optical character recognition (OCR). Work on OCR for low-resource Kazakh scr…