activity
20202026
most citedMEST: Accurate and Fast Memory-Economic Sparse Training Framework on the Edge

41 citations · 108 across the 32 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2025★ 3 cited

TSLA: A Task-Specific Learning Adaptation for Semantic Segmentation on Autonomous Vehicles Platform

Jun Liu, Zhenglun Kong, Pu Zhao +9

Autonomous driving platforms encounter diverse driving scenarios, each with varying hardware resources and precision requirements. Given the computational limitations of embedded d…

cs.CV2024

Fast and Memory-Efficient Video Diffusion Using Streamlined Inference

Zheng Zhan, Yushu Wu, Yifan Gong +7

The rapid progress in artificial intelligence-generated content (AIGC), especially with diffusion models, has significantly advanced development of high-quality video generation. H…

cs.CV2024

Exploring Token Pruning in Vision State Space Models

Zheng Zhan, Zhenglun Kong, Yifan Gong +8

State Space Models (SSMs) have the advantage of keeping linear computational complexity compared to attention modules in transformers, and have been applied to vision tasks as a ne…

cs.CV2023

GPU Accelerated Color Correction and Frame Warping for Real-time Video Stitching

Lu Yang, Zhenglun Kong, Ting Li +3

Traditional image stitching focuses on a single panorama frame without considering the spatial-temporal consistency in videos. The straightforward image stitching approach will cau…

cs.CV2022

Peeling the Onion: Hierarchical Reduction of Data Redundancy for Efficient Vision Transformer Training

Zhenglun Kong, Haoyu Ma, Geng Yuan +12

Vision transformers (ViTs) have recently obtained success in many applications, but their intensive computation and heavy memory usage at both training and inference time limit the…

cs.CV2022★ 1 cited

You Need Multiple Exiting: Dynamic Early Exiting for Accelerating Unified Vision Language Model

Shengkun Tang, Yaqing Wang, Zhenglun Kong +6

Large-scale Transformer models bring significant improvements for various downstream vision language tasks with a unified architecture. The performance improvements come with incre…