7 papers
PoInit-of-View: Poisoning Initialization of Views Transfers Across Multiple 3D Reconstruction Systems
Weijie Wang, Songlong Xing, Zhengyu Zhao +2
Poisoning input views of 3D reconstruction systems has been recently studied. However, we identify that existing studies simply backpropagate adversarial gradients through the 3D r…
TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation
Yan Shu, Bin Ren, Zhitong Xiong +4
Vision-language models (VLMs) have shown promise in earth observation (EO), yet they struggle with tasks that require grounding complex spatial reasoning in precise pixel-level vis…
Panoramic Affordance Prediction
Zixin Zhang, Chenfei Liao, Hongfei Zhang +10
Affordance prediction serves as a critical bridge between perception and action in embodied AI. However, existing research is confined to pinhole camera models, which suffer from n…
High-Fidelity 3D Facial Avatar Synthesis with Controllable Fine-Grained Expressions
Yikang He, Jichao Zhang, Wei Wang +2
Facial expression editing methods can be mainly categorized into two types based on their architectures: 2D-based and 3D-based methods. The former lacks 3D face modeling capabiliti…
DVD: Deterministic Video Depth Estimation with Generative Priors
Hongfei Zhang, Harold Haodong Chen, Chenfei Liao +12
Existing video depth estimation faces a fundamental trade-off: generative models suffer from stochastic geometric hallucinations and scale drift, while discriminative models demand…
Stable-Hair v2: Real-World Hair Transfer via Multiple-View Diffusion Model
Kuiyuan Sun, Yuxuan Zhang, Jichao Zhang +4
While diffusion-based methods have shown impressive capabilities in capturing diverse and complex hairstyles, their ability to generate consistent and high-quality multi-view outpu…