2 papers
cs.CV2025
MoniRefer: A Real-world Large-scale Multi-modal Dataset based on Roadside Infrastructure for 3D Visual Grounding
Panquan Yang, Junfei Huang, Zongzhangbao Yin +9
3D visual grounding aims to localize the object in 3D point cloud scenes that semantically corresponds to given natural language sentences. It is very critical for roadside infrast…
cs.CV2024
WEM-GAN: Wavelet transform based facial expression manipulation
Dongya Sun, Yunfei Hu, Xianzhe Zhang +1
Facial expression manipulation aims to change human facial expressions without affecting face recognition. In order to transform the facial expressions to target expressions, previ…