69 citations · 82 across the 18 of their papers we have counts for
18 papers
MVBoost: Boost 3D Reconstruction with Multi-View Refinement
Xiangyu Liu, Xiaomei Zhang, Zhiyuan Ma +2
Recent advancements in 3D object reconstruction have been remarkable, yet most current 3D models rely heavily on existing 3D datasets. The scarcity of diverse 3D datasets results i…
Revisiting Marr in Face: The Building of 2D--2.5D--3D Representations in Deep Neural Networks
Xiangyu Zhu, Chang Yu, Jiankuo Zhao +3
David Marr's seminal theory of vision proposes that the human visual system operates through a sequence of three stages, known as the 2D sketch, the 2.5D sketch, and the 3D model.…
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy
Te Yang, Jian Jia, Xiangyu Zhu +9
Large Language Models (LLMs) have strong instruction-following capability to interpret and execute tasks as directed by human commands. Multimodal Large Language Models (MLLMs) hav…
S2TD-Face: Reconstruct a Detailed 3D Face with Controllable Texture from a Single Sketch
Zidu Wang, Xiangyu Zhu, Jiang Yu +2
3D textured face reconstruction from sketches applicable in many scenarios such as animation, 3D avatars, artistic design, missing people search, etc., is a highly promising but un…
DiffSpeaker: Speech-Driven 3D Facial Animation with Diffusion Transformer
Zhiyuan Ma, Xiangyu Zhu, Guojun Qi +3
Speech-driven 3D facial animation is important for many multimedia applications. Recent work has shown promise in using either Diffusion models or Transformer architectures for thi…
Modeling Spoof Noise by De-spoofing Diffusion and its Application in Face Anti-spoofing
Bin Zhang, Xiangyu Zhu, Xiaoyu Zhang +1
Face anti-spoofing is crucial for ensuring the security and reliability of face recognition systems. Several existing face anti-spoofing methods utilize GAN-like networks to detect…