Publications (95)
Probabilistic Self-supervised Learning via Scoring Rules Minimization
Amirhossein Vahidi, Simon SchoÃer, Lisa Wimmer +4
In this paper, we propose a novel probabilistic self-supervised learning via Scoring Rule Minimization (ProSMIN), which leverages the power of probabilistic models to enhance repre…
MambaIRv2: Attentive State Space Restoration
Hang Guo, Yong Guo, Yaohua Zha +5
The Mamba-based image restoration backbones have recently demonstrated significant potential in balancing global reception and computational efficiency. However, the inherent causa…
When SAM2 Meets Video Camouflaged Object Segmentation: A Comprehensive Evaluation and Adaptation
Yuli Zhou, Guolei Sun, Yawei Li +3
This study investigates the application and performance of the Segment Anything Model 2 (SAM2) in the challenging task of video camouflaged object segmentation (VCOS). VCOS involve…
Revisiting Adaptive Rounding with Vectorized Reparameterization for LLM Quantization
Yuli Zhou, Qingxuan Chen, Luca Benini +2
Adaptive Rounding has emerged as an alternative to round-to-nearest (RTN) for post-training quantization by enabling cross-element error cancellation. Yet, dense and element-wise r…
TinyMyo: a Tiny Foundation Model for Flexible EMG Signal Processing at the Edge
Matteo Fasulo, Giusy Spacone, Thorir Mar Ingolfsson +3
Objective: Surface electromyography (EMG) is a non-invasive sensing modality widely used in biomechanics, rehabilitation, prosthetic control, and human-machine interfaces. Despite…
Bringing Masked Autoencoders Explicit Contrastive Properties for Point Cloud Self-Supervised Learning
Bin Ren, Guofeng Mei, Danda Pani Paudel +6
Contrastive learning (CL) for Vision Transformers (ViTs) in image domains has achieved performance comparable to CL for traditional convolutional backbones. However, in 3D point cl…
The Heterogeneity Hypothesis: Finding Layer-Wise Differentiated Network Architectures
Yawei Li, Wen Li, Martin Danelljan +4
In this paper, we tackle the problem of convolutional neural network design. Instead of focusing on the design of the overall architecture, we investigate a design space that is us…
WaveFormer: A Lightweight Transformer Model for sEMG-based Gesture Recognition
Yanlong Chen, Mattia Orlandi, Pierangelo Maria Rapa +3
Human-machine interaction, particularly in prosthetic and robotic control, has seen progress with gesture recognition via surface electromyographic (sEMG) signals.However, classify…
Reference-based Image Super-Resolution with Deformable Attention Transformer
Jiezhang Cao, Jingyun Liang, Kai Zhang +4
Reference-based image super-resolution (RefSR) aims to exploit auxiliary reference (Ref) images to super-resolve low-resolution (LR) images. Recently, RefSR has been attracting gre…
Unsupervised Real-world Image Super Resolution via Domain-distance Aware Training
Yunxuan Wei, Shuhang Gu, Yawei Li +1
These days, unsupervised super-resolution (SR) has been soaring due to its practical and promising potential in real scenarios. The philosophy of off-the-shelf approaches lies in t…
The Eleventh NTIRE 2026 Efficient Super-Resolution Challenge Report
Bin Ren, Hang Guo, Yan Shu +60
This paper reviews the NTIRE 2026 challenge on efficient single-image super-resolution with a focus on the proposed solutions and results. The aim of this challenge is to devise a…
PhysioWave: A Multi-Scale Wavelet-Transformer for Physiological Signal Representation
Yanlong Chen, Mattia Orlandi, Pierangelo Maria Rapa +3
Physiological signals are often corrupted by motion artifacts, baseline drift, and other low-SNR disturbances, which pose significant challenges for analysis. Additionally, these s…
Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs
Yanlong Chen, Amirhossein Habibian, Luca Benini +1
Vision-Language Models (VLMs) achieve strong multimodal performance but are costly to deploy, and post-training quantization often causes significant accuracy loss. Despite its pot…
AIM 2020 Challenge on Efficient Super-Resolution: Methods and Results
Kai Zhang, Martin Danelljan, Yawei Li +75
This paper reviews the AIM 2020 challenge on efficient single image super-resolution with focus on the proposed solutions and results. The challenge task was to super-resolve an in…
The Tenth NTIRE 2025 Image Denoising Challenge Report
Lei Sun, Hang Guo, Bin Ren +91
This paper presents an overview of the NTIRE 2025 Image Denoising Challenge (Ï = 50), highlighting the proposed methodologies and corresponding results. The primary objective is t…
CiaoSR: Continuous Implicit Attention-in-Attention Network for Arbitrary-Scale Image Super-Resolution
Jiezhang Cao, Qin Wang, Yongqin Xian +7
Learning continuous image representations is recently gaining popularity for image super-resolution (SR) because of its ability to reconstruct high-resolution images with arbitrary…
FEMBA: Efficient and Scalable EEG Analysis with a Bidirectional Mamba Foundation Model
Anna Tegon, Thorir Mar Ingolfsson, Xiaying Wang +2
Accurate and efficient electroencephalography (EEG) analysis is essential for detecting seizures and artifacts in long-term monitoring, with applications spanning hospital diagnost…
Retina : Low-Power Eye Tracking with Event Camera and Spiking Hardware
Pietro Bonazzi, Sizhen Bian, Giovanni Lippolis +3
This paper introduces a neuromorphic methodology for eye tracking, harnessing pure event data captured by a Dynamic Vision Sensor (DVS) camera. The framework integrates a directly…
MaizeField3D: A Curated 3D Point Cloud and Procedural Model Dataset of Field-Grown Maize from a Diversity Panel
Elvis Kimara, Mozhgan Hadadi, Jackson Godbersen +6
The development of artificial intelligence (AI) and machine learning (ML) based tools for 3D phenotyping, especially for maize, has been limited due to the lack of large and divers…
DHP: Differentiable Meta Pruning via HyperNetworks
Yawei Li, Shuhang Gu, Kai Zhang +2
Network pruning has been the driving force for the acceleration of neural networks and the alleviation of model storage/transmission burden. With the advent of AutoML and neural ar…
Joint Visual Denoising and Classification using Deep Learning
Gang Chen, Yawei Li, Sargur N. Srihari
Visual restoration and recognition are traditionally addressed in pipeline fashion, i.e. denoising followed by classification. Instead, observing correlations between the two tasks…
Unified Embedding Alignment for Open-Vocabulary Video Instance Segmentation
Hao Fang, Peng Wu, Yawei Li +2
Open-Vocabulary Video Instance Segmentation (VIS) is attracting increasing attention due to its ability to segment and track arbitrary objects. However, the recent Open-Vocabulary…
Towards Efficient Graph Convolutional Networks for Point Cloud Handling
Yawei Li, He Chen, Zhaopeng Cui +4
In this paper, we aim at improving the computational efficiency of graph convolutional networks (GCNs) for learning on point clouds. The basic graph convolution that is typically c…
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators
Chi Zhang, Luca Colagrande, Renzo Andri +6
Multi-Head Attention (MHA) is a critical computational kernel in transformer-based AI models. Emerging scalable tile-based accelerator architectures integrate increasing numbers of…
Revisiting Random Channel Pruning for Neural Network Compression
Yawei Li, Kamil Adamczewski, Wen Li +3
Channel (or 3D filter) pruning serves as an effective way to accelerate the inference of neural networks. There has been a flurry of algorithms that try to solve this practical pro…
Efficient and Explicit Modelling of Image Hierarchies for Image Restoration
Yawei Li, Yuchen Fan, Xiaoyu Xiang +4
The aim of this paper is to propose a mechanism to efficiently and explicitly model image hierarchies in the global, regional, and local range for image restoration. To achieve tha…
NTIRE 2022 Challenge on Efficient Super-Resolution: Methods and Results
Yawei Li, Kai Zhang, Radu Timofte +108
This paper reviews the NTIRE 2022 challenge on efficient single image super-resolution with focus on the proposed solutions and results. The task of the challenge was to super-reso…
Transcending the Limit of Local Window: Advanced Super-Resolution Transformer with Adaptive Token Dictionary
Leheng Zhang, Yawei Li, Xingyu Zhou +2
Single Image Super-Resolution is a classic computer vision problem that involves estimating high-resolution (HR) images from low-resolution (LR) ones. Although deep neural networks…
Towards smart canopies: Algorithmic design of maize canopy architectures that maximize light use efficiency
Nasla Saleem, Talukder Zaki Jubery, Yan Zhou +4
We present a computational framework that integrates functional-structural plant modeling (FSPM) with an evolutionary algorithm to optimize three-dimensional maize canopy architect…
Group Sparsity: The Hinge Between Filter Pruning and Decomposition for Network Compression
Yawei Li, Shuhang Gu, Christoph Mayer +2
In this paper, we analyze two popular network compression techniques, i.e. filter pruning and low-rank decomposition, in a unified sense. By simply changing the way the sparsity re…
FastVAR: Linear Visual Autoregressive Modeling via Cached Token Pruning
Hang Guo, Yawei Li, Taolin Zhang +4
Visual Autoregressive (VAR) modeling has gained popularity for its shift towards next-scale prediction. However, existing VAR paradigms process the entire token map at each scale s…
LUNA: Efficient and Topology-Agnostic Foundation Model for EEG Signal Analysis
Berkay Döner, Thorir Mar Ingolfsson, Luca Benini +1
Electroencephalography (EEG) offers a non-invasive lens into human brain activity, but building large-scale models is hampered by topological heterogeneity: each public EEG data de…
Shapley Pruning for Neural Network Compression
Kamil Adamczewski, Yawei Li, Luc van Gool
Neural network pruning is a rich field with a variety of approaches. In this work, we propose to connect the existing pruning concepts such as leave-one-out pruning and oracle prun…
CEReBrO: Compact Encoder for Representations of Brain Oscillations Using Efficient Alternating Attention
Alexandru Dimofte, Glenn Anta Bucagu, Thorir Mar Ingolfsson +4
Electroencephalograph (EEG) is a crucial tool for studying brain activity. Recently, self-supervised learning methods leveraging large unlabeled datasets have emerged as a potentia…
FEMBA on the Edge: Physiologically-Aware Pre-Training, Quantization, and Deployment of a Bidirectional Mamba EEG Foundation Model on an Ultra-low Power Microcontroller
Anna Tegon, Nicholas Lehmann, Yawei Li +3
Objective: To enable continuous, long-term neuro-monitoring on wearable devices by overcoming the computational bottlenecks of Transformer-based Electroencephalography (EEG) founda…
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Bowen Zhang, Congchao Guo, Geng Yang +17
We introduce MiniMax-Speech, an autoregressive Transformer-based Text-to-Speech (TTS) model that generates high-quality speech. A key innovation is our learnable speaker encoder, w…
SPEAR: A System for Post-Quantization Error-Adaptive Recovery Enabling Efficient Low-Bit LLM Serving
Hongyuan Liu, Yawei Li, Zhiqiang Que +3
Efficient large language model (LLM) serving is increasingly constrained by deployment cost. Quantization is a key technique for reducing serving cost, yet even state-of-the-art 4-…
RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache
Junkai Zhang, Hang Guo, Luca Benini +1
Large language models (LLMs) have shown strong performance across diverse tasks, but their inference with long input contexts is bottlenecked by memory size and bandwidth. The Key-…
MarkushGrapher: Joint Visual and Textual Recognition of Markush Structures
Lucas Morin, Valéry Weber, Ahmed Nassar +4
The automated analysis of chemical literature holds promise to accelerate discovery in fields such as material science and drug development. In particular, search capabilities for…
LuMamba: Latent Unified Mamba for Electrode Topology-Invariant and Efficient EEG Modeling
Danaé Broustail, Anna Tegon, Thorir Mar Ingolfsson +2
Electroencephalography (EEG) enables non-invasive monitoring of brain activity across clinical and neurotechnology applications, yet building foundation models for EEG remains chal…
Efficient Autoregressive Video Diffusion with Dummy Head
Hang Guo, Zhaoyang Jia, Jiahao Li +5
The autoregressive video diffusion model has recently gained considerable research interest due to its causal modeling and iterative denoising. In this work, we identify that the m…
Accessing the Effect of Phyllotaxy and Planting Density on Light Use Efficiency in Field-Grown Maize using 3D Reconstructions
Nasla Saleem, Talukder Zaki Jubery, Aditya Balu +5
High-density planting is a widely adopted strategy to enhance maize productivity, yet it introduces challenges such as increased interplant competition and shading, which can limit…
The Ninth NTIRE 2024 Efficient Super-Resolution Challenge Report
Bin Ren, Yawei Li, Nancy Mehta +129
This paper provides a comprehensive review of the NTIRE 2024 challenge, focusing on efficient single-image super-resolution (ESR) solutions and their outcomes. The task of this cha…
Giant Negative Thermal Expansion Induced by the Synergistic Effects of Ferroelectrostriction and Spin-Crossover in PbTiO3-Based Perovskites
Zhao Pan, Jun Chen, Runze Yu +18
The discovery of unusual negative thermal expansion (NTE) provides the opportunity to control the common but much desired property of thermal expansion, which is valuable not only…
CamSAM2: Segment Anything Accurately in Camouflaged Videos
Yuli Zhou, Yawei Li, Yuqian Fu +3
Video camouflaged object segmentation (VCOS), aiming at segmenting camouflaged objects that seamlessly blend into their environment, is a fundamental vision task with various real-…
Dynamic Parameter Memory: Temporary LoRA-Enhanced LLM for Long-Sequence Emotion Recognition in Conversation
Jialong Mai, Xiaofen Xing, Yawei Li +4
Recent research has focused on applying speech large language model (SLLM) to improve speech emotion recognition (SER). However, the inherently high frame rate in speech modality s…
PanLUNA: An Efficient and Robust Query-Unified Multimodal Model for Edge Biosignal Intelligence
Marija Zelic, Anna Tegon, Yawei Li +2
Physiological foundation models (FMs) have shown promise for biosignal representation learning, yet most remain confined to a single modality such as EEG, ECG, or PPG, largely beca…
Sharing Key Semantics in Transformer Makes Efficient Image Restoration
Bin Ren, Yawei Li, Jingyun Liang +6
Image Restoration (IR), a classic low-level vision task, has witnessed significant advancements through deep models that effectively model global information. Notably, the emergenc…
VRT: A Video Restoration Transformer
Jingyun Liang, Jiezhang Cao, Yuchen Fan +5
Video restoration (e.g., video super-resolution) aims to restore high-quality frames from low-quality frames. Different from single image restoration, video restoration generally r…
FinerCut: Finer-grained Interpretable Layer Pruning for Large Language Models
Yang Zhang, Yawei Li, Xinpeng Wang +5
Overparametrized transformer networks are the state-of-the-art architecture for Large Language Models (LLMs). However, such models contain billions of parameters making large compu…
A Dual-Perspective Approach to Evaluating Feature Attribution Methods
Yawei Li, Yang Zhang, Kenji Kawaguchi +3
Feature attribution methods attempt to explain neural network predictions by identifying relevant features. However, establishing a cohesive framework for assessing feature attribu…
Fine-Grained Neural Network Explanation by Identifying Input Features with Predictive Information
Yang Zhang, Ashkan Khakzar, Yawei Li +3
One principal approach for illuminating a black-box neural network is feature attribution, i.e. identifying the importance of input features for the network's prediction. The predi…
One Shot vs. Iterative: Rethinking Pruning Strategies for Model Compression
MikoÅaj Janusz, Tomasz Wojnar, Yawei Li +2
Pruning is a core technique for compressing neural networks to improve computational efficiency. This process is typically approached in two ways: one-shot pruning, which involves…
Predicting charge density distribution of materials using a local-environment-based graph convolutional network
Sheng Gong, Tian Xie, Taishan Zhu +4
Electron charge density distribution of materials is one of the key quantities in computational materials science as theoretically it determines the ground state energy and practic…
Tera-MIND: Tera-scale mouse brain simulation via spatial mRNA-guided diffusion
Jiqing Wu, Ingrid Berg, Yawei Li +2
Holistic 3D modeling of molecularly defined brain structures is crucial for understanding complex brain functions. Using emerging tissue profiling technologies, researchers charted…
The Tenth NTIRE 2025 Efficient Super-Resolution Challenge Report
Bin Ren, Hang Guo, Lei Sun +143
This paper presents a comprehensive review of the NTIRE 2025 Challenge on Single-Image Efficient Super-Resolution (ESR). The challenge aimed to advance the development of deep mode…
TinyTracker: Ultra-Fast and Ultra-Low-Power Edge Vision In-Sensor for Gaze Estimation
Pietro Bonazzi, Thomas Ruegg, Sizhen Bian +2
Intelligent edge vision tasks encounter the critical challenge of ensuring power and latency efficiency due to the typically heavy computational load they impose on edge platforms.…
Beyond GSD-as-Token: Continuous Scale Conditioning for Remote Sensing VLMs
Song Zhang, Yanlong Chen, Yilin Li +4
Remote sensing vision-language models (RS-VLMs) face a fundamental mismatch with natural-image counterparts: the same geographic object exhibits radically different visual evidence…
Robotic Grasping and Placement Controlled by EEG-Based Hybrid Visual and Motor Imagery
Yichang Liu, Tianyu Wang, Ziyi Ye +4
We present a framework that integrates EEG-based visual and motor imagery (VI/MI) with robotic control to enable real-time, intention-driven grasping and placement. Motivated by th…
Towards Extreme Pruning of LLMs with Plug-and-Play Mixed Sparsity
Chi Xu, Gefei Zhang, Yantong Zhu +4
N:M structured pruning is essential for large language models (LLMs) because it can remove less important network weights and reduce the memory and computation requirements. Existi…
Explaining COVID-19 and Thoracic Pathology Model Predictions by Identifying Informative Input Features
Ashkan Khakzar, Yang Zhang, Wejdene Mansour +5
Neural networks have demonstrated remarkable performance in classification and regression tasks on chest X-rays. In order to establish trust in the clinical routine, the networks'…
AttributionLab: Faithfulness of Feature Attribution Under Controllable Environments
Yang Zhang, Yawei Li, Hannah Brown +5
Feature attribution explains neural network outputs by identifying relevant input features. The attribution has to be faithful, meaning that the attributed features must mirror the…
Cluster, Split, Fuse, and Update: Meta-Learning for Open Compound Domain Adaptive Semantic Segmentation
Rui Gong, Yuhua Chen, Danda Pani Paudel +5
Open compound domain adaptation (OCDA) is a domain adaptation setting, where target domain is modeled as a compound of multiple unknown homogeneous domains, which brings the advant…
Empowering Image Recovery_ A Multi-Attention Approach
Juan Wen, Yawei Li, Chao Zhang +3
We propose Diverse Restormer (DART), a novel image restoration method that effectively integrates information from various sources (long sequences, local and global regions, featur…
ATD: Improved Transformer with Adaptive Token Dictionary for Image Restoration
Leheng Zhang, Wei Long, Yawei Li +3
Recently, Transformers have gained significant popularity in image restoration tasks such as image super-resolution and denoising, owing to their superior performance. However, bal…
Fractal-IR: A Unified Framework for Efficient and Scalable Image Restoration
Yawei Li, Bin Ren, Jingyun Liang +5
While vision transformers achieve significant breakthroughs in various image restoration (IR) tasks, it is still challenging to efficiently scale them across multiple types of degr…
Restore Anything Model via Efficient Degradation Adaptation
Bin Ren, Eduard Zamfir, Zongwei Wu +6
With the proliferation of mobile devices, the need for an efficient model to restore any degraded image has become increasingly significant and impactful. Traditional approaches ty…
Machine Learning Applications in Lung Cancer Diagnosis, Treatment and Prognosis
Yawei Li, Xin Wu, Ping Yang +2
The recent development of imaging and sequencing technologies enables systematic advances in the clinical study of lung cancer. Meanwhile, the human mind is limited in effectively…
3D Appearance Super-Resolution with Deep Learning
Yawei Li, Vagia Tsiminaki, Radu Timofte +2
We tackle the problem of retrieving high-resolution (HR) texture maps of objects that are captured from multiple view points. In the multi-view case, model-based super-resolution (…
Q-MambaIR: Accurate Quantized Mamba for Efficient Image Restoration
Yujie Chen, Haotong Qin, Zhang Zhang +3
State-Space Models (SSMs) have attracted considerable attention in Image Restoration (IR) due to their ability to scale linearly sequence length while effectively capturing long-di…
Analyzing the Effects of Handling Data Imbalance on Learned Features from Medical Images by Looking Into the Models
Ashkan Khakzar, Yawei Li, Yang Zhang +5
One challenging property lurking in medical datasets is the imbalanced data distribution, where the frequency of the samples between the different classes is not balanced. Training…
IntLoRA: Integral Low-rank Adaptation of Quantized Diffusion Models
Hang Guo, Yawei Li, Tao Dai +2
Fine-tuning pre-trained diffusion models under limited budgets has gained great success. In particular, the recent advances that directly fine-tune the quantized weights using Low-…
Q-Segment: Segmenting Images In-Sensor for Vessel-Based Medical Diagnosis
Pietro Bonazzi, Yawei Li, Sizhen Bian +1
This paper addresses the growing interest in deploying deep learning models directly in-sensor. We present "Q-Segment", a quantized real-time segmentation algorithm, and conduct a…
Video Super-Resolution Transformer
Jiezhang Cao, Yawei Li, Kai Zhang +1
Video super-resolution (VSR), with the aim to restore a high-resolution video from its corresponding low-resolution version, is a spatial-temporal sequence prediction problem. Rece…
The Third Challenge on Image Denoising at NTIRE 2026: Methods and Results
Lei Sun, Hang Guo, Bin Ren +6
This paper reports on the NTIRE 2026 Challenge on Image Denoising, specifically focusing on the high-noise regime (). The competition investigates advanced neural architect…
Any Image Restoration via Efficient Spatial-Frequency Degradation Adaptation
Bin Ren, Eduard Zamfir, Zongwei Wu +7
Restoring multiple degradations efficiently via just one model has become increasingly significant and impactful, especially with the proliferation of mobile devices. Traditional s…
Calibrating LLMs with Information-Theoretic Evidential Deep Learning
Yawei Li, David Rügamer, Bernd Bischl +1
Fine-tuned large language models (LLMs) often exhibit overconfidence, particularly when trained on small datasets, resulting in poor calibration and inaccurate uncertainty estimate…
Plug-and-Play Image Restoration with Deep Denoiser Prior
Kai Zhang, Yawei Li, Wangmeng Zuo +3
Recent works on plug-and-play image restoration have shown that a denoiser can implicitly serve as the image prior for model-based methods to solve many inverse problems. Such a pr…
S-CEReBrO: Breaking the Memory Barrier in Continuous EEG Monitoring
Glenn Anta Bucagu, Thorir Mar Ingolfsson, Yawei Li +1
The paper introduces S-CEReBrO, a streaming Transformer architecture that uses a windowed alternating attention mechanism to keep memory usage constant during continuous EEG monito…
Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs
Hang Guo, Yawei Li, Luca Benini
Recent advances in Large Language Model (LLM) compression, such as quantization and pruning, have achieved notable success. However, as these techniques gradually approach their re…
HiM2SAM: Enhancing SAM2 with Hierarchical Motion Estimation and Memory Optimization towards Long-term Tracking
Ruixiang Chen, Guolei Sun, Yawei Li +2
This paper presents enhancements to the SAM2 framework for video object tracking task, addressing challenges such as occlusions, background clutter, and target reappearance. We int…
Procedural Generation of 3D Maize Plant Architecture from LIDAR Data
Mozhgan Hadadi, Mehdi Saraeian, Jackson Godbersen +8
This study introduces a robust framework for generating procedural 3D models of maize (Zea mays) plants from LiDAR point cloud data, offering a scalable alternative to traditional…
A Compute&Memory Efficient Model-Driven Neural 5G Receiver for Edge AI-assisted RAN
Mahdi Abdollahpour, Marco Bertuletti, Yichao Zhang +3
Artificial intelligence approaches for base-band processing for radio receivers have demonstrated significant performance gains. Most of the proposed methods are characterized by h…
Steering Large Reasoning Models towards Concise Reasoning via Flow Matching
Yawei Li, Benjamin Bergner, Yinghan Zhao +3
Large Reasoning Models (LRMs) excel at complex reasoning tasks, but their efficiency is often hampered by overly verbose outputs. Prior steering methods attempt to address this iss…
SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models
Wei Huang, Haotong Qin, Yangdong Liu +7
Post-training quantization (PTQ) is an effective technique for compressing large language models (LLMs). However, while uniform-precision quantization is computationally efficient,…
Key-Graph Transformer for Image Restoration
Bin Ren, Yawei Li, Jingyun Liang +5
While it is crucial to capture global information for effective image restoration (IR), integrating such cues into transformer-based methods becomes computationally expensive, espe…
Finetuning and Quantization of EEG-Based Foundational BioSignal Models on ECG and PPG Data for Blood Pressure Estimation
Bálint Tóth, Dominik Senti, Thorir Mar Ingolfsson +4
Blood pressure (BP) is a key indicator of cardiovascular health. As hypertension remains a global cause of morbidity and mortality, accurate, continuous, and non-invasive BP monito…
Efficient Degradation-agnostic Image Restoration via Channel-Wise Functional Decomposition and Manifold Regularization
Bin Ren, Yawei Li, Xu Zheng +6
Degradation-agnostic image restoration aims to handle diverse corruptions with one unified model, but faces fundamental challenges in balancing efficiency and performance across di…
Practical Blind Image Denoising via Swin-Conv-UNet and Data Synthesis
Kai Zhang, Yawei Li, Jingyun Liang +6
While recent years have witnessed a dramatic upsurge of exploiting deep neural networks toward solving image denoising, existing methods mostly rely on simple noise assumptions, su…
LocalViT: Analyzing Locality in Vision Transformers
Yawei Li, Kai Zhang, Jiezhang Cao +4
The aim of this paper is to study the influence of locality mechanisms in vision transformers. Transformers originated from machine translation and are particularly good at modelli…
Learning Filter Basis for Convolutional Neural Network Compression
Yawei Li, Shuhang Gu, Luc Van Gool +1
Convolutional neural networks (CNNs) based solutions have achieved state-of-the-art performances for many computer vision tasks, including classification and super-resolution of im…
Ultra-Efficient On-Device Object Detection on AI-Integrated Smart Glasses with TinyissimoYOLO
Julian Moosmann, Pietro Bonazzi, Yawei Li +4
Smart glasses are rapidly gaining advanced functions thanks to cutting-edge computing technologies, especially accelerated hardware architectures, and tiny Artificial Intelligence…
Word Recognition with Deep Conditional Random Fields
Gang Chen, Yawei Li, Sargur N. Srihari
Recognition of handwritten words continues to be an important problem in document analysis and recognition. Existing approaches extract hand-engineered features from word images--w…
PIRM Challenge on Perceptual Image Enhancement on Smartphones: Report
Andrey Ignatov, Radu Timofte, Thang Van Vu +45
This paper reviews the first challenge on efficient perceptual image enhancement with the focus on deploying deep learning models on smartphones. The challenge consisted of two tra…
Hierarchical Information Flow for Generalized Efficient Image Restoration
Yawei Li, Bin Ren, Jingyun Liang +5
While vision transformers show promise in numerous image restoration (IR) tasks, the challenge remains in efficiently generalizing and scaling up a model for multiple IR tasks. To…