From the 1 of 7 linked papers with an AI index.
7 papers
JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
Shawn Li, Wei Yang, Jike Zhong +11
The paper introduces JigShape, a benchmark of interlocking jigsaw puzzles designed to test visual‑geometric reasoning in vision‑language models, and shows that current zero‑shot an…
Learning to Reason by Analogy via Retrieval-Augmented Reinforcement Fine-Tuning
Zilin Xiao, Qi Ma, Chun-cheng Jason Chen +4
Retrieval-augmented generation (RAG) has become a standard mechanism for grounding language models in external knowledge, yet conventional retrieval based on lexical or semantic si…
EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation
Ruozhen He, Meng Wei, Ziyan Yang +1
Multi-shot video generation extends single-shot generation to coherent visual narratives, yet maintaining consistent characters, objects, and locations across shots remains a chall…
Beyond Referring Expressions: Scenario Comprehension Visual Grounding
Ruozhen He, Nisarg A. Shah, Qihua Dong +3
Existing visual grounding benchmarks primarily evaluate alignment between image regions and literal referring expressions, where models can often succeed by matching a prominent na…
Learning from Synthetic Data for Visual Grounding
Ruozhen He, Ziyan Yang, Paola Cascante-Bonilla +2
This paper extensively investigates the effectiveness of synthetic training data to improve the capabilities of vision-and-language models for grounding textual descriptions to ima…
Grounding Language Models for Visual Entity Recognition
Zilin Xiao, Ming Gong, Paola Cascante-Bonilla +3
We introduce AutoVER, an Autoregressive model for Visual Entity Recognition. Our model extends an autoregressive Multi-modal Large Language Model by employing retrieval augmented c…