10 papers
Composing Concepts from Images and Videos via Concept-prompt Binding
Xianghao Kong, Zeyu Zhang, Yuwei Guo +3
Visual concept composition, which aims to integrate different elements from images and videos into a single, coherent visual output, still falls short in accurately extracting comp…
Decoupling Defense Strategies for Robust Image Watermarking
Jiahui Chen, Zehang Deng, Zeyu Zhang +3
Deep learning-based image watermarking, while robust against conventional distortions, remains vulnerable to advanced adversarial and regeneration attacks. Conventional countermeas…
PartRAG: Retrieval-Augmented Part-Level 3D Generation and Editing
Peize Li, Zeyu Zhang, Hao Tang
Single-image 3D generation with part-level structure remains challenging: learned priors struggle to cover the long tail of part geometries and maintain multi-view consistency, and…
StereoAdapter-2: Globally Structure-Consistent Underwater Stereo Depth Estimation
Zeyu Ren, Xiang Li, Yiran Wang +2
Stereo depth estimation is fundamental to underwater robotic perception, yet suffers from severe domain shifts caused by wavelength-dependent light attenuation, scattering, and ref…
HiVid: LLM-Guided Video Saliency For Content-Aware VOD And Live Streaming
Jiahui Chen, Bo Peng, Lianchen Jia +3
Content-aware streaming requires dynamic, chunk-level importance weights to optimize subjective quality of experience (QoE). However, direct human annotation is prohibitively expen…
MatE: Material Extraction from Single-Image via Geometric Prior
Zeyu Zhang, Wei Zhai, Jian Yang +1
The creation of high-fidelity, physically-based rendering (PBR) materials remains a bottleneck in many graphics pipelines, typically requiring specialized equipment and expert-driv…