activity
20232026
most citedMixFormerV2: Efficient Fully Transformer Tracking

28 citations · 28 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV2026

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

Guozhen Zhang, Xuerui Qiu, Yutao Cui +11

Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space. In this paper, we present HYDR…

cs.CV2025

Differentiable Solver Search for Fast Diffusion Sampling

Shuai Wang, Zexian Li, Qipeng zhang +5

Diffusion models have demonstrated remarkable generation quality but at the cost of numerous function evaluations. Recently, advanced ODE-based solvers have been developed to mitig…

cs.CV2025

DMM: Building a Versatile Image Generation Model via Distillation-Based Model Merging

Tianhui Song, Weixin Feng, Shuai Wang +4

The success of text-to-image (T2I) generation models has spurred a proliferation of numerous model checkpoints fine-tuned from the same base model on various specialized datasets.…

cs.CV2024

FlowDCN: Exploring DCN-like Architectures for Fast Image Generation with Arbitrary Resolution

Shuai Wang, Zexian Li, Tianhui Song +4

Arbitrary-resolution image generation still remains a challenging task in AIGC, as it requires handling varying resolutions and aspect ratios while maintaining high visual quality.…

cs.CV2024

Accelerating Image Generation with Sub-path Linear Approximation Model

Chen Xu, Tianhui Song, Weixin Feng +4

Diffusion models have significantly advanced the state of the art in image, audio, and video generation tasks. However, their applications in practical scenarios are hindered by sl…

cs.CV2023★ 28 cited

MixFormerV2: Efficient Fully Transformer Tracking

Yutao Cui, Tianhui Song, Gangshan Wu +1

Transformer-based trackers have achieved strong accuracy on the standard benchmarks. However, their efficiency remains an obstacle to practical deployment on both GPU and CPU platf…