activity
20242026
most citedResCLIP: Residual Attention for Training-free Dense Vision-language Inference

10 citations · 10 across the 3 of their papers we have counts for

collaborators

10 papers

cs.CV2026

SA-Homo: Scale Adaptive Homography Estimation for Scale Variation Scenarios

Shangxuan Xie, Haifeng Wu, Yuhang Wang +2

Homography estimation, as one of the fundamental problems in computer vision, remains challenged by scale variation scenarios where image pairs potentially exhibit significant scal…

cs.CV202610 cited

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference

Yuhang Yang, Jinhong Deng, Wen Li +1

While vision-language models like CLIP have shown remarkable success in open-vocabulary tasks, their application is currently confined to image-level tasks, and they still struggle…

cs.CL2026

MixRea: Benchmarking Explicit-Implicit Reasoning in Large Language Models

Yuanqing Cai, Ziyi Huang, Minhao Liu +3

Large language models (LLMs) are increasingly integrated into high-stakes decision-making. Inspired by the theory of \emph{inattentional blindness} in human cognition, we investiga…

cs.CV2026

Tuning-Free Adaptive Style Incorporation for Structure-Consistent Text-Driven Style Transfer

Yanqi Ge, Jiaqi Liu, Qingnan Fan +6

In this work, we target the task of text-driven style transfer in the context of text-to-image (T2I) diffusion models. The main challenge is consistent structure preservation while…

cs.CV2025

The Devil is in Attention Sharing: Improving Complex Non-rigid Image Editing Faithfulness via Attention Synergy

Zhuo Chen, Fanyue Wei, Runze Xu +4

Training-free image editing with large diffusion models has become practical, yet faithfully performing complex non-rigid edits (e.g., pose or shape changes) remains highly challen…

cs.CV2025

OmniMotion-X: Versatile Multimodal Whole-Body Motion Generation

Guowei Xu, Yuxuan Bian, Ailing Zeng +6

This paper introduces OmniMotion-X, a versatile multimodal framework for whole-body human motion generation, leveraging an autoregressive diffusion transformer in a unified sequenc…