8 papers
DriveCtrl: Conditioned Sim-to-Real Driving Video Generation
Haonan Zhao, Yiting Wang, Jingkun Chen +3
Large-scale labelled driving video data is essential for training autonomous driving systems. Although simulation offers scalable and fully annotated data, the domain gap between s…
AuralSAM2: Enabling SAM2 Hear Through Pyramid Audio-Visual Feature Prompting
Yuyuan Liu, Yuanhong Chen, Chong Wang +6
Segment Anything Model 2 (SAM2) exhibits strong generalisation for promptable segmentation in video clips; however, its integration with the audio modality remains underexplored. E…
Reinforcing 3D Understanding in Point-VLMs via Geometric Reward Credit Assignment
Jingkun Chen, Ruoshi Xu, Mingqi Gao +2
Point-Vision-Language Models promise to empower embodied agents with executable spatial reasoning, yet they frequently succumb to geometric hallucination where predicted 3D structu…
Evo: Autoregressive-Diffusion Large Language Models with Evolving Balance
Junde Wu, Minhao Hu, Jiayuan Zhu +7
We introduce \textbf{Evo}, a duality latent trajectory model that bridges autoregressive (AR) and diffusion-based language generation within a continuous evolutionary generative fr…
TRACE: Temporally Reliable Anatomically-Conditioned 3D CT Generation with Enhanced Efficiency
Minye Shao, Xingyu Miao, Haoran Duan +7
3D medical image generation is essential for data augmentation and patient privacy, calling for reliable and efficient models suited for clinical practice. However, current methods…
VDNeRF: Vision-only Dynamic Neural Radiance Field for Urban Scenes
Zhengyu Zou, Jingfeng Li, Hao Li +5
Neural Radiance Fields (NeRFs) implicitly model continuous three-dimensional scenes using a set of images with known camera poses, enabling the rendering of photorealistic novel vi…