6 papers
Structure-Level Disentangled Diffusion for Few-Shot Chinese Font Generation
Jie Li, Suorong Yang, Jian Zhao +1
Few-shot Chinese font generation aims to synthesize new characters in a target style using only a handful of reference images. Achieving accurate content rendering and faithful sty…
Vectra: A New Metric, Dataset, and Model for Visual Quality Assessment in E-Commerce In-Image Machine Translation
Qingyu Wu, Yuxuan Han, Haijun Li +5
In-Image Machine Translation (IIMT) powers cross-border e-commerce product listings; existing research focuses on machine translation evaluation, while visual rendering quality is…
CVAM-Pose: Conditional Variational Autoencoder for Multi-Object Monocular Pose Estimation
Jianyu Zhao, Wei Quan, Bogdan J. Matuszewski
Estimating rigid objects' poses is one of the fundamental problems in computer vision, with a range of applications across automation and augmented reality. Most existing approache…
CoSIGN: Few-Step Guidance of ConSIstency Model to Solve General INverse Problems
Jiankun Zhao, Bowen Song, Liyue Shen
Diffusion models have been demonstrated as strong priors for solving general inverse problems. Most existing Diffusion model-based Inverse Problem Solvers (DIS) employ a plug-and-p…
Technique Report of CVPR 2024 PBDL Challenges
Ying Fu, Yu Li, Shaodi You +96
The intersection of physics-based vision and deep learning presents an exciting frontier for advancing computer vision technologies. By leveraging the principles of physics to info…
The SkatingVerse Workshop & Challenge: Methods and Results
Jian Zhao, Lei Jin, Jianshu Li +16
The SkatingVerse Workshop & Challenge aims to encourage research in developing novel and accurate methods for human action understanding. The SkatingVerse dataset used for the Skat…