3 papers
cs.CV2026
GaussianFusion: Unified 3D Gaussian Representation for Multi-Modal Fusion Perception
Xiao Zhao, Chang Liu, Mingxu Zhu +5
The bird's-eye view (BEV) representation enables multi-sensor features to be fused within a unified space, serving as the primary approach for achieving comprehensive 3D perception…
cs.CV2026
Video-HOCA: A Diagnostic Benchmark for Physical Anomaly Reasoning in Video-LLMs
Chang Liu, Yunfan Ye, Qingyang Zhou +5
We introduce Video-HOCA, a diagnostic benchmark for physical anomaly reasoning in videos. Video-HOCA uses an Ontological-Causal taxonomy to distinguish violations of an entity's ow…
cs.CV2026
Decoupling Vision and Language: Codebook Anchored Visual Adaptation
Jason Wu, Tianchen Zhao, Chang Liu +7
Large Vision-Language Models (LVLMs) use their vision encoders to translate images into representations for downstream reasoning, but the encoders often underperform in domain-spec…