6 papers
Training-Free Representation Guidance for Diffusion Models with a Representation Alignment Projector
Wenqiang Zu, Shenghao Xie, Bo Lei +1
Recent progress in generative modeling has enabled high-quality visual synthesis with diffusion-based frameworks, supporting controllable sampling and large-scale training. Inferen…
Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
Zhiyuan Jiang, Shenghao Xie, Wenyi Li +8
Grounding is a fundamental capability for building graphical user interface (GUI) agents. Although existing approaches rely on large-scale bounding box supervision, they still face…
NANO3D: A Training-Free Approach for Efficient 3D Editing Without Masks
Junliang Ye, Shenghao Xie, Ruowen Zhao +5
3D object editing is essential for interactive content creation in gaming, animation, and robotics, yet current approaches remain inefficient, inconsistent, and often fail to prese…
Exploring Representation Invariance in Finetuning
Wenqiang Zu, Shenghao Xie, Hao Chen +9
Foundation models pretrained on large-scale natural images are widely adapted to various cross-domain low-resource downstream tasks, benefiting from generalizable and transferable…
Embedded Visual Prompt Tuning
Wenqiang Zu, Shenghao Xie, Qing Zhao +2
Foundation models pre-trained on large-scale data have been widely witnessed to achieve success in various natural imaging downstream tasks. Parameter-efficient fine-tuning (PEFT)…
Pre-trained Models Succeed in Medical Imaging with Representation Similarity Degradation
Wenqiang Zu, Shenghao Xie, Hao Chen +1
This paper investigates the critical problem of representation similarity evolution during cross-domain transfer learning, with particular focus on understanding why pre-trained mo…