4 citations · 9 across the 3 of their papers we have counts for
8 papers
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection
Shuyu Wang, Weiqi Li, Qian Wang +2
Recent advances in AI-generated content (AIGC) have significantly accelerated image editing techniques, driving increasing demand for diverse and fine-grained edits. Despite these…
HiCoM: Hierarchical Coherent Motion for Streamable Dynamic Scene with 3D Gaussian Splatting
Qiankun Gao, Jiarui Meng, Chengxiang Wen +2
The online reconstruction of dynamic scenes from multi-view streaming videos faces significant challenges in training, rendering and storage efficiency. Harnessing superior learnin…
CPA: Camera-pose-awareness Diffusion Transformer for Video Generation
Yuelei Wang, Jian Zhang, Pengtao Jiang +3
Despite the significant advancements made by Diffusion Transformer (DiT)-based methods in video generation, there remains a notable gap with controllable camera pose perspectives.…
Large Spatial Model: End-to-end Unposed Images to Semantic 3D
Zhiwen Fan, Jian Zhang, Wenyan Cong +10
Reconstructing and understanding 3D structures from a limited number of images is a well-established problem in computer vision. Traditional methods usually break this task into mu…
CRoSS: Diffusion Model Makes Controllable, Robust and Secure Image Steganography
Jiwen Yu, Xuanyu Zhang, Youmin Xu +1
Current image steganography techniques are mainly focused on cover-based methods, which commonly have the risk of leaking secret images and poor robustness against degraded contain…
The Foreseeable Future: Self-Supervised Learning to Predict Dynamic Scenes for Indoor Navigation
Hugues Thomas, Jian Zhang, Timothy D. Barfoot
We present a method for generating, predicting, and using Spatiotemporal Occupancy Grid Maps (SOGM), which embed future semantic information of real dynamic scenes. We present an a…