2 papers
cs.CV2026
VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning
Hengbo Xu, Shengjie Jin, Yanbiao Ma +1
With the rapid advancement of large multimodal models (LMMs), inference-time overhead has become a key bottleneck for real-world deployment. Existing methods typically prune visual…
cs.CV2024
MonoMAE: Enhancing Monocular 3D Detection through Depth-Aware Masked Autoencoders
Xueying Jiang, Sheng Jin, Xiaoqin Zhang +2
Monocular 3D object detection aims for precise 3D localization and identification of objects from a single-view image. Despite its recent progress, it often struggles while handlin…