activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2025

Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models

Jonathan Lee, Xingrui Wang, Jiawei Peng +9

We propose Perceptual Taxonomy, a structured process of scene understanding that first recognizes objects and their spatial configurations, then infers task-relevant properties suc…

cs.CV2025

EvoWorld: Evolving Panoramic World Generation with Explicit 3D Memory

Jiahao Wang, Luoxin Ye, TaiMing Lu +8

Humans possess a remarkable ability to mentally explore and replay 3D environments they have previously experienced. Inspired by this mental process, we present EvoWorld: a world m…

cs.CV2025

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models

Wufei Ma, Luoxin Ye, Celso M de Melo +2

Humans naturally understand 3D spatial relationships, enabling complex reasoning like predicting collisions of vehicles from different directions. Current large multimodal models (…

cs.CV2025

GenEx: Generating an Explorable World

Taiming Lu, Tianmin Shu, Junfei Xiao +8

Understanding, navigating, and exploring the 3D physical real world has long been a central challenge in the development of artificial intelligence. In this work, we take a step to…

cs.CV2024

Efficient Large Multi-modal Models via Visual Context Compression

Jieneng Chen, Luoxin Ye, Ju He +3

While significant advancements have been made in compressed representations for text embeddings in large language models (LLMs), the compression of visual tokens in multi-modal LLM…