activity
20222026
most citedRank-DETR for High Quality Object Detection

25 citations · 41 across the 18 of their papers we have counts for

collaborators
Showing cs.CVShow all

22 papers · 1 filter

cs.CV2026

MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale

Zhicong Tang, Zhao Zhang, Jingye Chen +6

Layered image generation and editing is a fundamental capability that enables layer-wise reuse, editing, and composition of generated visual content, analogous to word-level editin…

cs.CV2026

Bridging the RGB-IR Gap: Consensus and Discrepancy Modeling for Text-Guided Multispectral Detection

Jiaqi Wu, Zhen Wang, Enhao Huang +5

Text-guided multispectral object detection uses text semantics to guide semantic-aware cross-modal interaction between RGB and IR for more robust perception. However, notable limit…

cs.CV2026

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling

Keming Wu, Zuhao Yang, Kaichen Zhang +24

Recent visual generation models have made major progress in photorealism, typography, instruction following, and interactive editing, yet they still struggle with spatial reasoning…

cs.CV2025

Few-Step Distillation for Text-to-Image Generation: A Practical Guide

Yifan Pu, Yizeng Han, Zhiwei Tang +4

Diffusion distillation has dramatically accelerated class-conditional image synthesis, but its applicability to open-ended text-to-image (T2I) generation is still unclear. We prese…

cs.CV2025

Linear Differential Vision Transformer: Learning Visual Contrasts via Pairwise Differentials

Yifan Pu, Jixuan Ying, Qixiu Li +7

Vision Transformers (ViTs) have become a universal backbone for both image recognition and image generation. Yet their Multi-Head Self-Attention (MHSA) layer still performs a quadr…

cs.CV2025

Emulating Human-like Adaptive Vision for Efficient and Flexible Machine Visual Perception

Yulin Wang, Yang Yue, Huanqian Wang +11

Human vision is highly adaptive, efficiently sampling intricate environments by sequentially fixating on task-relevant regions. In contrast, prevailing machine vision models passiv…