2 citations · 2 across the 3 of their papers we have counts for
4 papers · 1 filter
Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends
Jiuming Liu, Chaojun Ni, Mengmeng Liu +7
With rapid development of large language models and diffusion-based content generation, world modeling has attracted increasing research attention, benefiting various downstream do…
Neural-MMGS: Multi-modal Neural Gaussian Splats for Large-Scale Scene Reconstruction
Sitian Shen, Georgi Pramatarov, Yifu Tao +1
This paper proposes Neural-MMGS, a novel neural 3DGS framework for multimodal large-scale scene reconstruction that fuses multiple sensing modalities in a per-gaussian compact, lea…
DiffCLIP: Leveraging Stable Diffusion for Language Grounded 3D Classification
Sitian Shen, Zilin Zhu, Linqian Fan +2
Large pre-trained models have had a significant impact on computer vision by enabling multi-modal learning, where the CLIP model has achieved impressive results in image classifica…
Increasing Visual Awareness in Multimodal Neural Machine Translation from an Information Theoretic Perspective
Baijun Ji, Tong Zhang, Yicheng Zou +2
Multimodal machine translation (MMT) aims to improve translation quality by equipping the source sentence with its corresponding image. Despite the promising performance, MMT model…