activity
20232026
most citedOccWorld: Learning a 3D Occupancy World Model for Autonomous Driving

5 citations · 6 across the 30 of their papers we have counts for

collaborators
Showing cs.CVShow all

50 papers · 1 filter

cs.CV2026

SM4RT: Learning Structured Motion Geometry for 4D Reconstruction

Shing Ho J. Lin, Wenzhao Zheng, Dong Zhuo +3

Geometry Foundation Models (GFMs) have substantially advanced monocular 3D reconstruction, yet extending this capability to 4D dynamic understanding remains a fundamental challenge…

cs.CV2026

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution

Xumin Yu, Zuyan Liu, Zhenyu Yang +5

A unified representation for text and vision is a natural pursuit, as it enables simpler multimodal modeling and more efficient training. However, representing images as discrete s…

cs.CV2026

Point Cloud Diffusion with Global and Local Reconstruction for Instance-Level 3D Anomaly Detection

Linchun Wu, Qin Zou, Jiwen Lu +1

3D anomaly detection in point clouds is critical for high-precision industrial manufacturing. Reconstruction-based methods have laid a strong foundation by detecting 3D anomalies t…

cs.CV2026

Segment Anything with Motion, Geometry, and Semantic Adaptation for Complex Nonlinear Visual Object Tracking

Deyi Zhu, Yuji Wang, Yong Liu +4

Traditional visual object tracking (VOT) methods typically rely on task-specific supervised training, limiting their generalization to unseen objects and challenging scenarios with…

cs.CV2026

IVGT: Implicit Visual Geometry Transformer for Neural Scene Representation

Yuqi Wu, Tianyu Hu, Wenzhao Zheng +4

Reconstructing coherent 3D geometry and appearance from unposed multi-view images is a fundamental yet challenging problem in computer vision. Most existing visual geometry foundat…

cs.CV2026

BAMI: Training-Free Bias Mitigation in GUI Grounding

Borui Zhang, Bo Zhang, Bo Wang +6

GUI grounding is a critical capability for enabling GUI agents to execute tasks such as clicking and dragging. However, in complex scenarios like the ScreenSpot-Pro benchmark, exis…