most citedDefining and Evaluating Visual Language Models' Basic Spatial Abilities: A Perspective from Psychometrics

2 citations · 2 across the 9 of their papers we have counts for

collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2026

SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models

Sirun Li, Minghao Liu, Ling Dai +4

Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitra…

cs.CV2026

ForceForget: Reinforcement Concept Removal for Enhancing Safety in Text-to-Image Models

Dong Han, Yong Li

With the advance of generative AI, the text-to-image (T2I) model has the ability to generate various contents. However, T2I models still can generate unsafe contents. To alleviate…

cs.CV2026

HDINO: A Concise and Efficient Open-Vocabulary Detector

Hao Zhang, Yiqun Wang, Qinran Lin +2

Despite the growing interest in open-vocabulary object detection in recent years, most existing methods rely heavily on manually curated fine-grained training datasets as well as r…

cs.CV2025

EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent

Jiaao Li, Kaiyuan Li, Chen Gao +2

Egomotion videos are first-person recordings where the view changes continuously due to the agent's movement. As they serve as the primary visual input for embodied AI agents, maki…

cs.CV2025

DIMM: Decoupled Multi-hierarchy Kalman Filter for 3D Object Tracking

Jirong Zha, Yuxuan Fan, Kai Li +4

State estimation is challenging for 3D object tracking with high maneuverability, as the target's state transition function changes rapidly, irregularly, and is unknown to the esti…

cs.CV2025

Balanced Token Pruning: Accelerating Vision Language Models Beyond Local Optimization

Kaiyuan Li, Xiaoyue Chen, Chen Gao +2

Large Vision-Language Models (LVLMs) have shown impressive performance across multi-modal tasks by encoding images into thousands of tokens. However, the large number of image toke…