8 papers
Initialize to Generalize: A Stronger Initialization Pipeline for Sparse-View 3DGS
Feng Zhou, Wenkai Guo, Pu Cao +2
Sparse-view 3D Gaussian Splatting (3DGS) often overfits to the training views, leading to artifacts like blurring in novel view rendering. Prior work addresses it either by enhanci…
Learning a Unified Policy for Position and Force Control in Legged Loco-Manipulation
Peiyuan Zhi, Peiyang Li, Jianqin Yin +2
Robotic loco-manipulation tasks often involve contact-rich interactions with the environment, requiring the joint modeling of contact force and robot position. However, recent visu…
BiTAA: A Bi-Task Adversarial Attack for Object Detection and Depth Estimation via 3D Gaussian Splatting
Yixun Zhang, Feng Zhou, Jianqin Yin
Camera-based perception is critical to autonomous driving yet remains vulnerable to task-specific adversarial manipulations in object detection and monocular depth estimation. Most…
L2HCount:Generalizing Crowd Counting from Low to High Crowd Density via Density Simulation
Guoliang Xu, Jianqin Yin, Ren Zhang +3
Since COVID-19, crowd-counting tasks have gained wide applications. While supervised methods are reliable, annotation is more challenging in high-density scenes due to small head s…
Exploring Position Encoding in Diffusion U-Net for Training-free High-resolution Image Generation
Feng Zhou, Pu Cao, Yiyang Ma +2
Denoising higher-resolution latents via a pre-trained U-Net leads to repetitive and disordered image patterns. Although recent studies make efforts to improve generative quality by…
InstruGen: Automatic Instruction Generation for Vision-and-Language Navigation Via Large Multimodal Models
Yu Yan, Rongtao Xu, Jiazhao Zhang +3
Recent research on Vision-and-Language Navigation (VLN) indicates that agents suffer from poor generalization in unseen environments due to the lack of realistic training environme…