3 papers
cs.CV2023
Contrast-augmented Diffusion Model with Fine-grained Sequence Alignment for Markup-to-Image Generation
Guojin Zhong, Jin Yuan, Pan Wang +3
The recently rising markup-to-image generation poses greater challenges as compared to natural image generation, due to its low tolerance for errors as well as the complex sequence…
cs.CV2023
PVPUFormer: Probabilistic Visual Prompt Unified Transformer for Interactive Image Segmentation
Xu Zhang, Kailun Yang, Jiacheng Lin +3
Integration of diverse visual prompts like clicks, scribbles, and boxes in interactive image segmentation significantly facilitates users' interaction as well as improves interacti…
cs.CV2023
SSD-MonoDETR: Supervised Scale-aware Deformable Transformer for Monocular 3D Object Detection
Xuan He, Fan Yang, Kailun Yang +5
Transformer-based methods have demonstrated superior performance for monocular 3D object detection recently, which aims at predicting 3D attributes from a single 2D image. Most exi…