10 citations · 10 across the 3 of their papers we have counts for
10 papers
SA-Homo: Scale Adaptive Homography Estimation for Scale Variation Scenarios
Shangxuan Xie, Haifeng Wu, Yuhang Wang +2
Homography estimation, as one of the fundamental problems in computer vision, remains challenged by scale variation scenarios where image pairs potentially exhibit significant scal…
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference
Yuhang Yang, Jinhong Deng, Wen Li +1
While vision-language models like CLIP have shown remarkable success in open-vocabulary tasks, their application is currently confined to image-level tasks, and they still struggle…
MixRea: Benchmarking Explicit-Implicit Reasoning in Large Language Models
Yuanqing Cai, Ziyi Huang, Minhao Liu +3
Large language models (LLMs) are increasingly integrated into high-stakes decision-making. Inspired by the theory of \emph{inattentional blindness} in human cognition, we investiga…
Tuning-Free Adaptive Style Incorporation for Structure-Consistent Text-Driven Style Transfer
Yanqi Ge, Jiaqi Liu, Qingnan Fan +6
In this work, we target the task of text-driven style transfer in the context of text-to-image (T2I) diffusion models. The main challenge is consistent structure preservation while…
The Devil is in Attention Sharing: Improving Complex Non-rigid Image Editing Faithfulness via Attention Synergy
Zhuo Chen, Fanyue Wei, Runze Xu +4
Training-free image editing with large diffusion models has become practical, yet faithfully performing complex non-rigid edits (e.g., pose or shape changes) remains highly challen…
OmniMotion-X: Versatile Multimodal Whole-Body Motion Generation
Guowei Xu, Yuxuan Bian, Ailing Zeng +6
This paper introduces OmniMotion-X, a versatile multimodal framework for whole-body human motion generation, leveraging an autoregressive diffusion transformer in a unified sequenc…