activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models

Chak-Wing Mak, Guanyu Zhu, Boyi Zhang +16

Modern foundational Multimodal Large Language Models (MLLMs) and video world models have advanced significantly in mathematical, common-sense, and visual reasoning, but their grasp…

cs.CV2025

AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs

Yuan Zhang, Chun-Kai Fan, Sicheng Yu +6

Inspired by text prompts in large language models, visual prompts have been explored to enhance the perceptual capabilities of large vision-language models (LVLMs). However, perfor…

cs.CV2024

EVA: An Embodied World Model for Future Video Anticipation

Xiaowei Chi, Chun-Kai Fan, Hengyuan Zhang +8

Video generation models have made significant progress in simulating future states, showcasing their potential as world simulators in embodied scenarios. However, existing models o…

cs.CV2024

SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Yuan Zhang, Chun-Kai Fan, Junpeng Ma +8

In vision-language models (VLMs), visual tokens usually bear a significant amount of computational overhead despite sparsity of information in them when compared to text tokens. To…

cs.CV2024

Unveiling the Tapestry of Consistency in Large Vision-Language Models

Yuan Zhang, Fei Xiao, Tao Huang +7

Large vision-language models (LVLMs) have recently achieved rapid progress, exhibiting great perception and reasoning abilities concerning visual information. However, when faced w…