8 papers
MoWorld: A Flash World Model
Team Moxin, Deyi Ji, Tianrun Chen +37
The future of World Models depends not only on scaling model capability, but also on scaling practicality and inference efficiency. High-frame-rate inference enables responsive per…
Distilling Diffusion Models to Efficient 3D LiDAR Scene Completion
Shengyuan Zhang, An Zhao, Ling Yang +7
Diffusion models have been applied to 3D LiDAR scene completion due to their strong training stability and high completion quality. However, the slow sampling speed limits the prac…
Syllables to Scenes: Literary-Guided Free-Viewpoint 3D Scene Synthesis from Japanese Haiku
Chunan Yu, Yidong Han, Chaotao Ding +6
In the era of the metaverse, where immersive technologies redefine human experiences, translating abstract literary concepts into navigable 3D environments presents a fundamental c…
Let Human Sketches Help: Empowering Challenging Image Segmentation Task with Freehand Sketches
Ying Zang, Runlong Cao, Jianqi Zhang +8
Sketches, with their expressive potential, allow humans to convey the essence of an object through even a rough contour. For the first time, we harness this expressive potential to…
Img2CAD: Conditioned 3D CAD Model Generation from Single Image with Structured Visual Geometry
Tianrun Chen, Chunan Yu, Yuanqi Hu +8
In this paper, we propose Img2CAD, the first approach to our knowledge that uses 2D image inputs to generate CAD models with editable parameters. Unlike existing AI methods for 3D…
SAM2-Adapter: Evaluating & Adapting Segment Anything 2 in Downstream Tasks: Camouflage, Shadow, Medical Image Segmentation, and More
Tianrun Chen, Ankang Lu, Lanyun Zhu +7
The advent of large models, also known as foundation models, has significantly transformed the AI research landscape, with models like Segment Anything (SAM) achieving notable succ…