Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
NoteVQA: Benchmarking VLMs on Real-Life Questions from Human Communities
Haonan Jiang, Guojian Zhan, Jiancong Xie +6
Vision-language models (VLMs) increasingly power consumer-facing AI search, yet evaluating them on the diversity of everyday visual questions remains challenging. Existing benchmar…
cs.AI2025
HardcoreLogic: Challenging Large Reasoning Models with Long-tail Logic Puzzle Games
Jingcong Liang, Shijun Wan, Xuehai Wu +5
Large Reasoning Models (LRMs) have demonstrated impressive performance on complex tasks, including logical puzzle games that require deriving solutions satisfying all constraints.…