1 citations · 1 across the 4 of their papers we have counts for
9 papers
SDMatte: Grafting Diffusion Models for Interactive Matting
Longfei Huang, Yu Liang, Hao Zhang +6
Recent interactive matting methods have shown satisfactory performance in capturing the primary regions of objects, but they fall short in extracting fine-grained details in edge r…
Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning
Xinbin Yuan, Jian Zhang, Kaixin Li +8
Graphical User Interface (GUI) agents have made substantial strides in understanding and executing user instructions across diverse platforms. Yet, grounding these instructions to…
Any-to-Bokeh: Arbitrary-Subject Video Refocusing with Video Diffusion Model
Yang Yang, Siming Zheng, Qirui Yang +6
Diffusion models have recently emerged as powerful tools for camera simulation, enabling both geometric transformations and realistic optical effects. Among these, image-based boke…
MagicTryOn: Harnessing Diffusion Transformer for Garment-Preserving Video Virtual Try-on
Guangyuan Li, Siming Zheng, Hao Zhang +6
Video Virtual Try-On (VVT) aims to synthesize garments that appear natural across consecutive video frames, capturing both their dynamics and interactions with human motion. Despit…
Photography Perspective Composition: Towards Aesthetic Perspective Recommendation
Lujian Yao, Siming Zheng, Xinbin Yuan +5
Traditional photography composition approaches are dominated by 2D cropping-based methods. However, these methods fall short when scenes contain poorly arranged subjects. Professio…
M2N2V2: Multi-Modal Unsupervised and Training-free Interactive Segmentation
Markus Karmann, Peng-Tao Jiang, Bo Li +1
We present Markov Map Nearest Neighbor V2 (M2N2V2), a novel and simple, yet effective approach which leverages depth guidance and attention maps for unsupervised and training-free…