3 papers
cs.CV2026
ReScene: Structured Indoor Scene Reconstruction from Multi-View Captures
Haoran Xu, Lechao Zhang, Daoguo Dong +2
Constructing simulation-ready 3D scenes from multi-view captures is a key bottleneck for Embodied Artificial Intelligence, as downstream tasks require object-level structure, expli…
cs.CV2026
Efficient Multimodal Large Language Models: A Survey
Yizhang Jin, Jian Li, Yexin Liu +10
In the past year, Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in tasks such as visual question answering, visual understanding and reasoning.…
cs.CV2025
MOS: Modeling Object-Scene Associations in Generalized Category Discovery
Zhengyuan Peng, Jinpeng Ma, Zhimin Sun +4
Generalized Category Discovery (GCD) is a classification task that aims to classify both base and novel classes in unlabeled images, using knowledge from a labeled dataset. In GCD,…