2 papers
cs.CV2026
V-ABS: Action-Observer Driven Beam Search for Dynamic Visual Reasoning
Zhiwei Ning, Xuanang Gao, Jiaxi Cao +6
Multimodal large language models (MLLMs) have achieved remarkable success in general perception, yet complex multi-step visual reasoning remains a persistent challenge. Although re…
cs.CV2026
Fore-Mamba3D: Mamba-based Foreground-Enhanced Encoding for 3D Object Detection
Zhiwei Ning, Xuanang Gao, Jiaxi Cao +5
Linear modeling methods like Mamba have been merged as the effective backbone for the 3D object detection task. However, previous Mamba-based methods utilize the bidirectional enco…