works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.CV2026

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

Shawn Li, Wei Yang, Jike Zhong +11

The paper introduces JigShape, a benchmark of interlocking jigsaw puzzles designed to test visual‑geometric reasoning in vision‑language models, and shows that current zero‑shot an…

cs.CL2026

Learning to Reason by Analogy via Retrieval-Augmented Reinforcement Fine-Tuning

Zilin Xiao, Qi Ma, Chun-cheng Jason Chen +4

Retrieval-augmented generation (RAG) has become a standard mechanism for grounding language models in external knowledge, yet conventional retrieval based on lexical or semantic si…

cs.CV2026

EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation

Ruozhen He, Meng Wei, Ziyan Yang +1

Multi-shot video generation extends single-shot generation to coherent visual narratives, yet maintaining consistent characters, objects, and locations across shots remains a chall…

cs.CV2026

Beyond Referring Expressions: Scenario Comprehension Visual Grounding

Ruozhen He, Nisarg A. Shah, Qihua Dong +3

Existing visual grounding benchmarks primarily evaluate alignment between image regions and literal referring expressions, where models can often succeed by matching a prominent na…

cs.CV2024

Learning from Synthetic Data for Visual Grounding

Ruozhen He, Ziyan Yang, Paola Cascante-Bonilla +2

This paper extensively investigates the effectiveness of synthetic training data to improve the capabilities of vision-and-language models for grounding textual descriptions to ima…

cs.CV2024

Grounding Language Models for Visual Entity Recognition

Zilin Xiao, Ming Gong, Paola Cascante-Bonilla +3

We introduce AutoVER, an Autoregressive model for Visual Entity Recognition. Our model extends an autoregressive Multi-modal Large Language Model by employing retrieval augmented c…