works on

From the 1 of 13 linked papers with an AI index.

collaborators

13 papers

cs.CV2026

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

Shawn Li, Wei Yang, Jike Zhong +11

The paper introduces JigShape, a benchmark of interlocking jigsaw puzzles designed to test visual‑geometric reasoning in vision‑language models, and shows that current zero‑shot an…

cs.AI2026

Agentic Context Learning with Self-Discovered Specification

Jike Zhong, Ming Li, Yuxiang Lai +8

Context learning is an emerging inference-time task where LLMs must learn and apply novel, task-specific knowledge from intricate contexts absent from pre-training; even frontier m…

cs.LG2026

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning

Jike Zhong, Yuxiang Lai, Ming Li +5

Theory of Mind (ToM) is a must-acquire skill for modern foundation model systems to operate effectively and safely in the real world. Recent works have explored honing ToM via post…

cs.AI2026

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents

Yuxiang Lai, Peng Xia, Haonian Ji +8

Interactive agent benchmarks face a tension between scalable construction and realistic workflow evaluation. Hand-authored tasks are expensive to extend and revise, while static pr…

cs.CV2026

Are Video Models Emerging as Zero-Shot Learners and Reasoners in Medical Imaging?

Yuxiang Lai, Jike Zhong, Ming Li +2

Recent advances in large generative models have shown that simple autoregressive formulations, when scaled appropriately, can exhibit strong zero-shot generalization across domains…

cs.CV2026

VRIQ: Benchmarking and Analyzing Visual-Reasoning IQ of VLMs

Tina Khezresmaeilzadeh, Jike Zhong, Konstantinos Psounis

Recent progress in Vision Language Models (VLMs) has raised the question of whether they can reliably perform nonverbal reasoning. To this end, we introduce VRIQ (Visual Reasoning…