papers

Publications (23)

eess.IV2020

Automatically Searching for U-Net Image Translator Architecture

Han Shu, Yunhe Wang

Image translators have been successfully applied to many important low level image processing tasks. However, classical network architecture of image translator like U-Net, is borr…

cs.CV2026

TinySAM 2: Extreme Memory Compression for Efficient Track Anything Model

Zhaoyuan Ding, Yijing Yang, Han Shu +1

Segment Anything Model 2 (SAM 2) serves as a core foundation model in the field of video segmentation. Building upon the original SAM model, it introduces a memory bank mechanism a…

cs.LG2025

Expert Merging: Model Merging with Unsupervised Expert Alignment and Importance-Guided Layer Chunking

Dengming Zhang, Xiaowen Ma, Zhenliang Ni +4

Model merging, which combines multiple domain-specialized experts into a single model, offers a practical path to endow Large Language Models (LLMs) and Multimodal Large Language M…

eess.SY2021

Control Reconfiguration of Dynamical Systems for Improved Performance via Reverse- and Forward-engineering

Han Shu, Xuan Zhang, Na Li +1

This paper presents a control reconfiguration approach to improve the performance of two classes of dynamical systems. Motivated by recent research on re-engineering cyber-physical…

cs.CV2021

Coarse-to-Fine Searching for Efficient Generative Adversarial Networks

Jiahao Wang, Han Shu, Weihao Xia +2

This paper studies the neural architecture search (NAS) problem for developing efficient generator networks. Compared with deep models for visual recognition tasks, generative adve…

cs.CV2019

Attribute Aware Pooling for Pedestrian Attribute Recognition

Kai Han, Yunhe Wang, Han Shu +3

This paper expands the strength of deep convolutional neural networks (CNNs) to the pedestrian attribute recognition problem by devising a novel attribute aware pooling algorithm.…

cs.CL2026

Thinking-while-speaking: A Controlled, Interleaved Reasoning Method for Real-Time Speech Generation

Xuan Du, Qiangyu Yan, Wenshuo Li +4

The thinking-while-speaking paradigm aims to make AI communication more human. A key challenge is maintaining fluent speech while performing deep reasoning. Our method, InterRS, ta…

eess.SY2024

Transmission Benefits and Cost Allocation under Ambiguity

Han Shu, Jacob Mays

Disputes over cost allocation can present a significant barrier to investment in shared infrastructure. While it may be desirable to allocate cost in a way that corresponds to expe…

cs.CV2026

VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm

Zhenkai Wu, Xiaowen Ma, Zhenliang Ni +4

Vision-language models (VLMs) excel at image understanding tasks, but the large number of visual tokens imposes significant computational costs, hindering deployment on mobile devi…

cs.LG2025

Ada-MoGE: Adaptive Mixture of Gaussian Expert Model for Time Series Forecasting

Zhenliang Ni, Xiaowen Ma, Zhenkai Wu +3

Multivariate time series forecasts are widely used, such as industrial, transportation and financial forecasts. However, the dominant frequencies in time series may shift with the…

cs.LG2025

Federated Koopman-Reservoir Learning for Large-Scale Multivariate Time-Series Anomaly Detection

Long Tan Le, Tung-Anh Nguyen, Han Shu +3

The proliferation of edge devices has dramatically increased the generation of multivariate time-series (MVTS) data, essential for applications from healthcare to smart cities. Suc…

q-fin.TR2022

Beyond capacity: contractual form in electricity reliability obligations

Han Shu, Jacob Mays

Liberalized electricity markets often include resource adequacy mechanisms that require consumers to contract with generation resources well in advance of real-time operations. Whi…

cs.CV2025

ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding

Jialiang Kang, Han Shu, Wenshuo Li +2

Speculative decoding is a widely adopted technique for accelerating inference in large language models (LLMs), yet its application to vision-language models (VLMs) remains underexp…

cs.CV2020

VEGA: Towards an End-to-End Configurable AutoML Pipeline

Bochao Wang, Hang Xu, Jiajin Zhang +21

Automated Machine Learning (AutoML) is an important industrial solution for automatic discovery and deployment of the machine learning models. However, designing an integrated Auto…

cs.CV2020

Distilling portable Generative Adversarial Networks for Image Translation

Hanting Chen, Yunhe Wang, Han Shu +5

Despite Generative Adversarial Networks (GANs) have been widely used in various image-to-image translation tasks, they can be hardly applied on mobile devices due to their heavy co…

cs.CV2026

SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models

Yaozhi Wen, Jialong Guo, Zhenliang Ni +2

While Vision-Language Models (VLMs) have demonstrated remarkable performance in processing and understanding both text and images, their large parameter sizes lead to significant c…

cs.LG2024

ExCP: Extreme LLM Checkpoint Compression via Weight-Momentum Joint Shrinking

Wenshuo Li, Xinghao Chen, Han Shu +2

Large language models (LLM) have recently attracted significant attention in the field of artificial intelligence. However, the training process of these models poses significant c…

cs.DC2025

GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference

Phuong Tran, Tzu-Hao Liu, Long Tan Le +6

Large language models (LLMs) have revolutionized natural language processing, yet their high computational demands pose significant challenges for real-time inference, especially i…

cs.CV2026

SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuation

Jialiang Kang, Han Shu, Wenshuo Li +2

Speculative Jacobi Decoding (SJD) offers a draft-model-free approach to accelerate autoregressive text-to-image synthesis. However, the high-entropy nature of visual generation yie…

cs.CV2025

TinySAM: Pushing the Envelope for Efficient Segment Anything Model

Han Shu, Wenshuo Li, Yehui Tang +5

Recently segment anything model (SAM) has shown powerful segmentation capability and has drawn great attention in computer vision fields. Massive following works have developed var…

cs.CV2020

Optical Flow Distillation: Towards Efficient and Stable Video Style Transfer

Xinghao Chen, Yiman Zhang, Yunhe Wang +3

Video style transfer techniques inspire many exciting applications on mobile devices. However, their efficiency and stability are still far from satisfactory. To boost the transfer…

cs.AI2024

REPO: mplicit Reward Pairwise Difference based Empirical Preference Optimization

Long Tan Le, Han Shu, Tung-Anh Nguyen +2

While astonishingly capable, large Language Models (LLM) can sometimes produce outputs that deviate from human expectations. Such deviations necessitate an alignment phase to preve…

cs.CV2019

Co-Evolutionary Compression for Unpaired Image Translation

Han Shu, Yunhe Wang, Xu Jia +5

Generative adversarial networks (GANs) have been successfully used for considerable computer vision tasks, especially the image-to-image translation. However, generators in these n…