Publications (23)
Automatically Searching for U-Net Image Translator Architecture
Han Shu, Yunhe Wang
Image translators have been successfully applied to many important low level image processing tasks. However, classical network architecture of image translator like U-Net, is borr…
TinySAM 2: Extreme Memory Compression for Efficient Track Anything Model
Zhaoyuan Ding, Yijing Yang, Han Shu +1
Segment Anything Model 2 (SAM 2) serves as a core foundation model in the field of video segmentation. Building upon the original SAM model, it introduces a memory bank mechanism a…
Expert Merging: Model Merging with Unsupervised Expert Alignment and Importance-Guided Layer Chunking
Dengming Zhang, Xiaowen Ma, Zhenliang Ni +4
Model merging, which combines multiple domain-specialized experts into a single model, offers a practical path to endow Large Language Models (LLMs) and Multimodal Large Language M…
Control Reconfiguration of Dynamical Systems for Improved Performance via Reverse- and Forward-engineering
Han Shu, Xuan Zhang, Na Li +1
This paper presents a control reconfiguration approach to improve the performance of two classes of dynamical systems. Motivated by recent research on re-engineering cyber-physical…
Coarse-to-Fine Searching for Efficient Generative Adversarial Networks
Jiahao Wang, Han Shu, Weihao Xia +2
This paper studies the neural architecture search (NAS) problem for developing efficient generator networks. Compared with deep models for visual recognition tasks, generative adve…
Attribute Aware Pooling for Pedestrian Attribute Recognition
Kai Han, Yunhe Wang, Han Shu +3
This paper expands the strength of deep convolutional neural networks (CNNs) to the pedestrian attribute recognition problem by devising a novel attribute aware pooling algorithm.…
Thinking-while-speaking: A Controlled, Interleaved Reasoning Method for Real-Time Speech Generation
Xuan Du, Qiangyu Yan, Wenshuo Li +4
The thinking-while-speaking paradigm aims to make AI communication more human. A key challenge is maintaining fluent speech while performing deep reasoning. Our method, InterRS, ta…
Transmission Benefits and Cost Allocation under Ambiguity
Han Shu, Jacob Mays
Disputes over cost allocation can present a significant barrier to investment in shared infrastructure. While it may be desirable to allocate cost in a way that corresponds to expe…
VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
Zhenkai Wu, Xiaowen Ma, Zhenliang Ni +4
Vision-language models (VLMs) excel at image understanding tasks, but the large number of visual tokens imposes significant computational costs, hindering deployment on mobile devi…
Ada-MoGE: Adaptive Mixture of Gaussian Expert Model for Time Series Forecasting
Zhenliang Ni, Xiaowen Ma, Zhenkai Wu +3
Multivariate time series forecasts are widely used, such as industrial, transportation and financial forecasts. However, the dominant frequencies in time series may shift with the…
Federated Koopman-Reservoir Learning for Large-Scale Multivariate Time-Series Anomaly Detection
Long Tan Le, Tung-Anh Nguyen, Han Shu +3
The proliferation of edge devices has dramatically increased the generation of multivariate time-series (MVTS) data, essential for applications from healthcare to smart cities. Suc…
Beyond capacity: contractual form in electricity reliability obligations
Han Shu, Jacob Mays
Liberalized electricity markets often include resource adequacy mechanisms that require consumers to contract with generation resources well in advance of real-time operations. Whi…
ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding
Jialiang Kang, Han Shu, Wenshuo Li +2
Speculative decoding is a widely adopted technique for accelerating inference in large language models (LLMs), yet its application to vision-language models (VLMs) remains underexp…
VEGA: Towards an End-to-End Configurable AutoML Pipeline
Bochao Wang, Hang Xu, Jiajin Zhang +21
Automated Machine Learning (AutoML) is an important industrial solution for automatic discovery and deployment of the machine learning models. However, designing an integrated Auto…
Distilling portable Generative Adversarial Networks for Image Translation
Hanting Chen, Yunhe Wang, Han Shu +5
Despite Generative Adversarial Networks (GANs) have been widely used in various image-to-image translation tasks, they can be hardly applied on mobile devices due to their heavy co…
SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models
Yaozhi Wen, Jialong Guo, Zhenliang Ni +2
While Vision-Language Models (VLMs) have demonstrated remarkable performance in processing and understanding both text and images, their large parameter sizes lead to significant c…
ExCP: Extreme LLM Checkpoint Compression via Weight-Momentum Joint Shrinking
Wenshuo Li, Xinghao Chen, Han Shu +2
Large language models (LLM) have recently attracted significant attention in the field of artificial intelligence. However, the training process of these models poses significant c…
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
Phuong Tran, Tzu-Hao Liu, Long Tan Le +6
Large language models (LLMs) have revolutionized natural language processing, yet their high computational demands pose significant challenges for real-time inference, especially i…
SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuation
Jialiang Kang, Han Shu, Wenshuo Li +2
Speculative Jacobi Decoding (SJD) offers a draft-model-free approach to accelerate autoregressive text-to-image synthesis. However, the high-entropy nature of visual generation yie…
TinySAM: Pushing the Envelope for Efficient Segment Anything Model
Han Shu, Wenshuo Li, Yehui Tang +5
Recently segment anything model (SAM) has shown powerful segmentation capability and has drawn great attention in computer vision fields. Massive following works have developed var…
Optical Flow Distillation: Towards Efficient and Stable Video Style Transfer
Xinghao Chen, Yiman Zhang, Yunhe Wang +3
Video style transfer techniques inspire many exciting applications on mobile devices. However, their efficiency and stability are still far from satisfactory. To boost the transfer…
REPO: mplicit Reward Pairwise Difference based Empirical Preference Optimization
Long Tan Le, Han Shu, Tung-Anh Nguyen +2
While astonishingly capable, large Language Models (LLM) can sometimes produce outputs that deviate from human expectations. Such deviations necessitate an alignment phase to preve…
Co-Evolutionary Compression for Unpaired Image Translation
Han Shu, Yunhe Wang, Xu Jia +5
Generative adversarial networks (GANs) have been successfully used for considerable computer vision tasks, especially the image-to-image translation. However, generators in these n…