2 papers
cs.CV2026
Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D Reconstruction
Haoyu Zhang, Zeyu Zhang, Zedong Zhou +2
Transformer-based 3D reconstruction has emerged as a powerful paradigm for recovering geometry and appearance from multi-view observations, offering strong performance across chall…
cs.CV2025
CL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question Answering
Tianyu Huai, Jie Zhou, Xingjiao Wu +4
Multimodal large language models (MLLMs) have garnered widespread attention from researchers due to their remarkable understanding and generation capabilities in visual language ta…