From the 1 of 8 linked papers with an AI index.
8 papers
JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
Shawn Li, Wei Yang, Jike Zhong +11
The paper introduces JigShape, a benchmark of interlocking jigsaw puzzles designed to test visual‑geometric reasoning in vision‑language models, and shows that current zero‑shot an…
Learning Goal-Oriented Vision-and-Language Navigation with Self-Improving Demonstrations at Scale
Songze Li, Zun Wang, Gengze Zhou +8
Goal-oriented vision-language navigation requires robust exploration capabilities for agents to navigate to specified goals in unknown environments without step-by-step instruction…
Prune-Then-Plan: Step-Level Calibration for Stable Frontier Exploration in Embodied Question Answering
Noah Frahm, Prakrut Patel, Yue Zhang +3
Large vision-language models (VLMs) have improved embodied question answering (EQA) agents by providing strong semantic priors for open-vocabulary reasoning. However, when used dir…
Error-Driven Scene Editing for 3D Grounding in Large Language Models
Yue Zhang, Zun Wang, Han Lin +5
Despite recent progress in 3D-LLMs, they remain limited in accurately grounding language to visual and spatial elements in 3D environments. This limitation stems in part from train…
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
Shoubin Yu, Difan Liu, Ziqiao Ma +5
Recent video diffusion models have enhanced video editing, but it remains challenging to handle instructional editing and diverse tasks (e.g., adding, removing, changing) within a…
4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any Time
Ziqiao Ma, Xuweiyi Chen, Shoubin Yu +10
Can we scale 4D pretraining to learn general space-time representations that reconstruct an object from a few views at some times to any view at any time? We provide an affirmative…