3 papers
cs.CV2026
Audio-Driven Talking Face Generation with Blink Embedding and Hash Grid Landmarks Encoding
Yuhui Zhang, Hui Yu, Wei Liang +1
Dynamic Neural Radiance Fields (NeRF) have demonstrated considerable success in generating high-fidelity 3D models of talking portraits. Despite significant advancements in the ren…
cs.CV2025
AttentionDrag: Exploiting Latent Correlation Knowledge in Pre-trained Diffusion Models for Image Editing
Biao Yang, Muqi Huang, Yuhui Zhang +8
Traditional point-based image editing methods rely on iterative latent optimization or geometric transformations, which are either inefficient in their processing or fail to captur…
cs.CV2025
A 2D Semantic-Aware Position Encoding for Vision Transformers
Xi Chen, Shiyang Zhou, Muqi Huang +9
Vision transformers have demonstrated significant advantages in computer vision tasks due to their ability to capture long-range dependencies and contextual relationships through s…