6 papers
ReScene: Structured Indoor Scene Reconstruction from Multi-View Captures
Haoran Xu, Lechao Zhang, Daoguo Dong +2
Constructing simulation-ready 3D scenes from multi-view captures is a key bottleneck for Embodied Artificial Intelligence, as downstream tasks require object-level structure, expli…
Efficient Multimodal Large Language Models: A Survey
Yizhang Jin, Jian Li, Yexin Liu +10
In the past year, Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in tasks such as visual question answering, visual understanding and reasoning.…
MOS: Modeling Object-Scene Associations in Generalized Category Discovery
Zhengyuan Peng, Jinpeng Ma, Zhimin Sun +4
Generalized Category Discovery (GCD) is a classification task that aims to classify both base and novel classes in unlabeled images, using knowledge from a labeled dataset. In GCD,…
Textual Decomposition Then Sub-motion-space Scattering for Open-Vocabulary Motion Generation
Ke Fan, Jiangning Zhang, Ran Yi +6
Text-to-motion generation is a crucial task in computer vision, which generates the target 3D motion by the given text. The existing annotated datasets are limited in scale, result…
LLaVA-VSD: Large Language-and-Vision Assistant for Visual Spatial Description
Yizhang Jin, Jian Li, Jiangning Zhang +7
Visual Spatial Description (VSD) aims to generate texts that describe the spatial relationships between objects within images. Traditional visual spatial relationship classificatio…
AttentionPainter: An Efficient and Adaptive Stroke Predictor for Scene Painting
Yizhe Tang, Yue Wang, Teng Hu +5
Stroke-based Rendering (SBR) aims to decompose an input image into a sequence of parameterized strokes, which can be rendered into a painting that resembles the input image. Recent…