7 citations · 12 across the 4 of their papers we have counts for
5 papers
Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs
Xueying Jiang, Wenhao Li, Quanhao Qian +4
3D localization in Multimodal Large Language Models (MLLMs), including 3D object detection and 3D visual grounding, is fundamentally limited by camera intrinsic ambiguity: the same…
STS-Mixer: Spatio-Temporal-Spectral Mixer for 4D Point Cloud Video Understanding
Wenhao Li, Xueying Jiang, Gongjie Zhang +3
4D point cloud videos capture rich spatial and temporal dynamics of scenes which possess unique values in various 4D understanding tasks. However, most existing methods work in the…
Unsupervised Domain Adaptive 3D Detection with Multi-Level Consistency
Zhipeng Luo, Zhongang Cai, Changqing Zhou +7
Deep learning-based 3D object detection has achieved unprecedented success with the advent of large-scale autonomous driving datasets. However, drastic performance degradation rema…
FBC-GAN: Diverse and Flexible Image Synthesis via Foreground-Background Composition
Kaiwen Cui, Gongjie Zhang, Fangneng Zhan +2
Generative Adversarial Networks (GANs) have become the de-facto standard in image synthesis. However, without considering the foreground-background decomposition, existing GANs ten…
Unbalanced Feature Transport for Exemplar-based Image Translation
Fangneng Zhan, Yingchen Yu, Kaiwen Cui +7
Despite the great success of GANs in images translation with different conditioned inputs such as semantic segmentation and edge maps, generating high-fidelity realistic images wit…