papers

Publications (95)

cs.LG2023

Probabilistic Self-supervised Learning via Scoring Rules Minimization

Amirhossein Vahidi, Simon Schoßer, Lisa Wimmer +4

In this paper, we propose a novel probabilistic self-supervised learning via Scoring Rule Minimization (ProSMIN), which leverages the power of probabilistic models to enhance repre…

eess.IV2025

MambaIRv2: Attentive State Space Restoration

Hang Guo, Yong Guo, Yaohua Zha +5

The Mamba-based image restoration backbones have recently demonstrated significant potential in balancing global reception and computational efficiency. However, the inherent causa…

cs.CV2025

When SAM2 Meets Video Camouflaged Object Segmentation: A Comprehensive Evaluation and Adaptation

Yuli Zhou, Guolei Sun, Yawei Li +3

This study investigates the application and performance of the Segment Anything Model 2 (SAM2) in the challenging task of video camouflaged object segmentation (VCOS). VCOS involve…

cs.LG2026

Revisiting Adaptive Rounding with Vectorized Reparameterization for LLM Quantization

Yuli Zhou, Qingxuan Chen, Luca Benini +2

Adaptive Rounding has emerged as an alternative to round-to-nearest (RTN) for post-training quantization by enabling cross-element error cancellation. Yet, dense and element-wise r…

eess.SP2026

TinyMyo: a Tiny Foundation Model for Flexible EMG Signal Processing at the Edge

Matteo Fasulo, Giusy Spacone, Thorir Mar Ingolfsson +3

Objective: Surface electromyography (EMG) is a non-invasive sensing modality widely used in biomechanics, rehabilitation, prosthetic control, and human-machine interfaces. Despite…

cs.CV2024

Bringing Masked Autoencoders Explicit Contrastive Properties for Point Cloud Self-Supervised Learning

Bin Ren, Guofeng Mei, Danda Pani Paudel +6

Contrastive learning (CL) for Vision Transformers (ViTs) in image domains has achieved performance comparable to CL for traditional convolutional backbones. However, in 3D point cl…

cs.CV2021

The Heterogeneity Hypothesis: Finding Layer-Wise Differentiated Network Architectures

Yawei Li, Wen Li, Martin Danelljan +4

In this paper, we tackle the problem of convolutional neural network design. Instead of focusing on the design of the overall architecture, we investigate a design space that is us…

cs.CV2025

WaveFormer: A Lightweight Transformer Model for sEMG-based Gesture Recognition

Yanlong Chen, Mattia Orlandi, Pierangelo Maria Rapa +3

Human-machine interaction, particularly in prosthetic and robotic control, has seen progress with gesture recognition via surface electromyographic (sEMG) signals.However, classify…

cs.CV2022

Reference-based Image Super-Resolution with Deformable Attention Transformer

Jiezhang Cao, Jingyun Liang, Kai Zhang +4

Reference-based image super-resolution (RefSR) aims to exploit auxiliary reference (Ref) images to super-resolve low-resolution (LR) images. Recently, RefSR has been attracting gre…

cs.CV2020

Unsupervised Real-world Image Super Resolution via Domain-distance Aware Training

Yunxuan Wei, Shuhang Gu, Yawei Li +1

These days, unsupervised super-resolution (SR) has been soaring due to its practical and promising potential in real scenarios. The philosophy of off-the-shelf approaches lies in t…

cs.CV2026

The Eleventh NTIRE 2026 Efficient Super-Resolution Challenge Report

Bin Ren, Hang Guo, Yan Shu +60

This paper reviews the NTIRE 2026 challenge on efficient single-image super-resolution with a focus on the proposed solutions and results. The aim of this challenge is to devise a…

cs.LG2025

PhysioWave: A Multi-Scale Wavelet-Transformer for Physiological Signal Representation

Yanlong Chen, Mattia Orlandi, Pierangelo Maria Rapa +3

Physiological signals are often corrupted by motion artifacts, baseline drift, and other low-SNR disturbances, which pose significant challenges for analysis. Additionally, these s…

cs.CV2026

Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs

Yanlong Chen, Amirhossein Habibian, Luca Benini +1

Vision-Language Models (VLMs) achieve strong multimodal performance but are costly to deploy, and post-training quantization often causes significant accuracy loss. Despite its pot…

eess.IV2020

AIM 2020 Challenge on Efficient Super-Resolution: Methods and Results

Kai Zhang, Martin Danelljan, Yawei Li +75

This paper reviews the AIM 2020 challenge on efficient single image super-resolution with focus on the proposed solutions and results. The challenge task was to super-resolve an in…

cs.CV2025

The Tenth NTIRE 2025 Image Denoising Challenge Report

Lei Sun, Hang Guo, Bin Ren +91

This paper presents an overview of the NTIRE 2025 Image Denoising Challenge (σ = 50), highlighting the proposed methodologies and corresponding results. The primary objective is t…

cs.CV2023

CiaoSR: Continuous Implicit Attention-in-Attention Network for Arbitrary-Scale Image Super-Resolution

Jiezhang Cao, Qin Wang, Yongqin Xian +7

Learning continuous image representations is recently gaining popularity for image super-resolution (SR) because of its ability to reconstruct high-resolution images with arbitrary…

cs.LG2025

FEMBA: Efficient and Scalable EEG Analysis with a Bidirectional Mamba Foundation Model

Anna Tegon, Thorir Mar Ingolfsson, Xiaying Wang +2

Accurate and efficient electroencephalography (EEG) analysis is essential for detecting seizures and artifacts in long-term monitoring, with applications spanning hospital diagnost…

cs.CV2024

Retina : Low-Power Eye Tracking with Event Camera and Spiking Hardware

Pietro Bonazzi, Sizhen Bian, Giovanni Lippolis +3

This paper introduces a neuromorphic methodology for eye tracking, harnessing pure event data captured by a Dynamic Vision Sensor (DVS) camera. The framework integrates a directly…

cs.CV2025

MaizeField3D: A Curated 3D Point Cloud and Procedural Model Dataset of Field-Grown Maize from a Diversity Panel

Elvis Kimara, Mozhgan Hadadi, Jackson Godbersen +6

The development of artificial intelligence (AI) and machine learning (ML) based tools for 3D phenotyping, especially for maize, has been limited due to the lack of large and divers…

cs.CV2020

DHP: Differentiable Meta Pruning via HyperNetworks

Yawei Li, Shuhang Gu, Kai Zhang +2

Network pruning has been the driving force for the acceleration of neural networks and the alleviation of model storage/transmission burden. With the advent of AutoML and neural ar…

cs.CV2016

Joint Visual Denoising and Classification using Deep Learning

Gang Chen, Yawei Li, Sargur N. Srihari

Visual restoration and recognition are traditionally addressed in pipeline fashion, i.e. denoising followed by classification. Instead, observing correlations between the two tasks…

cs.CV2024

Unified Embedding Alignment for Open-Vocabulary Video Instance Segmentation

Hao Fang, Peng Wu, Yawei Li +2

Open-Vocabulary Video Instance Segmentation (VIS) is attracting increasing attention due to its ability to segment and track arbitrary objects. However, the recent Open-Vocabulary…

cs.CV2021

Towards Efficient Graph Convolutional Networks for Point Cloud Handling

Yawei Li, He Chen, Zhaopeng Cui +4

In this paper, we aim at improving the computational efficiency of graph convolutional networks (GCNs) for learning on point clouds. The basic graph convolution that is typically c…

cs.AR2025

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators

Chi Zhang, Luca Colagrande, Renzo Andri +6

Multi-Head Attention (MHA) is a critical computational kernel in transformer-based AI models. Emerging scalable tile-based accelerator architectures integrate increasing numbers of…

cs.CV2022

Revisiting Random Channel Pruning for Neural Network Compression

Yawei Li, Kamil Adamczewski, Wen Li +3

Channel (or 3D filter) pruning serves as an effective way to accelerate the inference of neural networks. There has been a flurry of algorithms that try to solve this practical pro…

cs.CV2023

Efficient and Explicit Modelling of Image Hierarchies for Image Restoration

Yawei Li, Yuchen Fan, Xiaoyu Xiang +4

The aim of this paper is to propose a mechanism to efficiently and explicitly model image hierarchies in the global, regional, and local range for image restoration. To achieve tha…

cs.CV2022

NTIRE 2022 Challenge on Efficient Super-Resolution: Methods and Results

Yawei Li, Kai Zhang, Radu Timofte +108

This paper reviews the NTIRE 2022 challenge on efficient single image super-resolution with focus on the proposed solutions and results. The task of the challenge was to super-reso…

cs.CV2024

Transcending the Limit of Local Window: Advanced Super-Resolution Transformer with Adaptive Token Dictionary

Leheng Zhang, Yawei Li, Xingyu Zhou +2

Single Image Super-Resolution is a classic computer vision problem that involves estimating high-resolution (HR) images from low-resolution (LR) ones. Although deep neural networks…

q-bio.QM2025

Towards smart canopies: Algorithmic design of maize canopy architectures that maximize light use efficiency

Nasla Saleem, Talukder Zaki Jubery, Yan Zhou +4

We present a computational framework that integrates functional-structural plant modeling (FSPM) with an evolutionary algorithm to optimize three-dimensional maize canopy architect…

cs.CV2020

Group Sparsity: The Hinge Between Filter Pruning and Decomposition for Network Compression

Yawei Li, Shuhang Gu, Christoph Mayer +2

In this paper, we analyze two popular network compression techniques, i.e. filter pruning and low-rank decomposition, in a unified sense. By simply changing the way the sparsity re…

cs.CV2025

FastVAR: Linear Visual Autoregressive Modeling via Cached Token Pruning

Hang Guo, Yawei Li, Taolin Zhang +4

Visual Autoregressive (VAR) modeling has gained popularity for its shift towards next-scale prediction. However, existing VAR paradigms process the entire token map at each scale s…

cs.LG2025

LUNA: Efficient and Topology-Agnostic Foundation Model for EEG Signal Analysis

Berkay Döner, Thorir Mar Ingolfsson, Luca Benini +1

Electroencephalography (EEG) offers a non-invasive lens into human brain activity, but building large-scale models is hampered by topological heterogeneity: each public EEG data de…

cs.LG2024

Shapley Pruning for Neural Network Compression

Kamil Adamczewski, Yawei Li, Luc van Gool

Neural network pruning is a rich field with a variety of approaches. In this work, we propose to connect the existing pruning concepts such as leave-one-out pruning and oracle prun…

cs.LG2025

CEReBrO: Compact Encoder for Representations of Brain Oscillations Using Efficient Alternating Attention

Alexandru Dimofte, Glenn Anta Bucagu, Thorir Mar Ingolfsson +4

Electroencephalograph (EEG) is a crucial tool for studying brain activity. Recently, self-supervised learning methods leveraging large unlabeled datasets have emerged as a potentia…

eess.SP2026

FEMBA on the Edge: Physiologically-Aware Pre-Training, Quantization, and Deployment of a Bidirectional Mamba EEG Foundation Model on an Ultra-low Power Microcontroller

Anna Tegon, Nicholas Lehmann, Yawei Li +3

Objective: To enable continuous, long-term neuro-monitoring on wearable devices by overcoming the computational bottlenecks of Transformer-based Electroencephalography (EEG) founda…

eess.AS2025

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Bowen Zhang, Congchao Guo, Geng Yang +17

We introduce MiniMax-Speech, an autoregressive Transformer-based Text-to-Speech (TTS) model that generates high-quality speech. A key innovation is our learnable speaker encoder, w…

cs.AR2026

SPEAR: A System for Post-Quantization Error-Adaptive Recovery Enabling Efficient Low-Bit LLM Serving

Hongyuan Liu, Yawei Li, Zhiqiang Que +3

Efficient large language model (LLM) serving is increasingly constrained by deployment cost. Quantization is a key technique for reducing serving cost, yet even state-of-the-art 4-…

cs.LG2026

RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache

Junkai Zhang, Hang Guo, Luca Benini +1

Large language models (LLMs) have shown strong performance across diverse tasks, but their inference with long input contexts is bottlenecked by memory size and bandwidth. The Key-…

cs.CV2025

MarkushGrapher: Joint Visual and Textual Recognition of Markush Structures

Lucas Morin, Valéry Weber, Ahmed Nassar +4

The automated analysis of chemical literature holds promise to accelerate discovery in fields such as material science and drug development. In particular, search capabilities for…

cs.AI2026

LuMamba: Latent Unified Mamba for Electrode Topology-Invariant and Efficient EEG Modeling

Danaé Broustail, Anna Tegon, Thorir Mar Ingolfsson +2

Electroencephalography (EEG) enables non-invasive monitoring of brain activity across clinical and neurotechnology applications, yet building foundation models for EEG remains chal…

cs.CV2026

Efficient Autoregressive Video Diffusion with Dummy Head

Hang Guo, Zhaoyang Jia, Jiahao Li +5

The autoregressive video diffusion model has recently gained considerable research interest due to its causal modeling and iterative denoising. In this work, we identify that the m…

cs.CV2025

Accessing the Effect of Phyllotaxy and Planting Density on Light Use Efficiency in Field-Grown Maize using 3D Reconstructions

Nasla Saleem, Talukder Zaki Jubery, Aditya Balu +5

High-density planting is a widely adopted strategy to enhance maize productivity, yet it introduces challenges such as increased interplant competition and shading, which can limit…

cs.CV2024

The Ninth NTIRE 2024 Efficient Super-Resolution Challenge Report

Bin Ren, Yawei Li, Nancy Mehta +129

This paper provides a comprehensive review of the NTIRE 2024 challenge, focusing on efficient single-image super-resolution (ESR) solutions and their outcomes. The task of this cha…

cond-mat.mtrl-sci2019

Giant Negative Thermal Expansion Induced by the Synergistic Effects of Ferroelectrostriction and Spin-Crossover in PbTiO3-Based Perovskites

Zhao Pan, Jun Chen, Runze Yu +18

The discovery of unusual negative thermal expansion (NTE) provides the opportunity to control the common but much desired property of thermal expansion, which is valuable not only…

cs.CV2025

CamSAM2: Segment Anything Accurately in Camouflaged Videos

Yuli Zhou, Yawei Li, Yuqian Fu +3

Video camouflaged object segmentation (VCOS), aiming at segmenting camouflaged objects that seamlessly blend into their environment, is a fundamental vision task with various real-…

cs.CL2025

Dynamic Parameter Memory: Temporary LoRA-Enhanced LLM for Long-Sequence Emotion Recognition in Conversation

Jialong Mai, Xiaofen Xing, Yawei Li +4

Recent research has focused on applying speech large language model (SLLM) to improve speech emotion recognition (SER). However, the inherently high frame rate in speech modality s…

cs.AI2026

PanLUNA: An Efficient and Robust Query-Unified Multimodal Model for Edge Biosignal Intelligence

Marija Zelic, Anna Tegon, Yawei Li +2

Physiological foundation models (FMs) have shown promise for biosignal representation learning, yet most remain confined to a single modality such as EEG, ECG, or PPG, largely beca…

cs.CV2024

Sharing Key Semantics in Transformer Makes Efficient Image Restoration

Bin Ren, Yawei Li, Jingyun Liang +6

Image Restoration (IR), a classic low-level vision task, has witnessed significant advancements through deep models that effectively model global information. Notably, the emergenc…

cs.CV2022

VRT: A Video Restoration Transformer

Jingyun Liang, Jiezhang Cao, Yuchen Fan +5

Video restoration (e.g., video super-resolution) aims to restore high-quality frames from low-quality frames. Different from single image restoration, video restoration generally r…

cs.LG2024

FinerCut: Finer-grained Interpretable Layer Pruning for Large Language Models

Yang Zhang, Yawei Li, Xinpeng Wang +5

Overparametrized transformer networks are the state-of-the-art architecture for Large Language Models (LLMs). However, such models contain billions of parameters making large compu…

cs.LG2024

A Dual-Perspective Approach to Evaluating Feature Attribution Methods

Yawei Li, Yang Zhang, Kenji Kawaguchi +3

Feature attribution methods attempt to explain neural network predictions by identifying relevant features. However, establishing a cohesive framework for assessing feature attribu…

cs.LG2021

Fine-Grained Neural Network Explanation by Identifying Input Features with Predictive Information

Yang Zhang, Ashkan Khakzar, Yawei Li +3

One principal approach for illuminating a black-box neural network is feature attribution, i.e. identifying the importance of input features for the network's prediction. The predi…

cs.LG2026

One Shot vs. Iterative: Rethinking Pruning Strategies for Model Compression

Mikołaj Janusz, Tomasz Wojnar, Yawei Li +2

Pruning is a core technique for compressing neural networks to improve computational efficiency. This process is typically approached in two ways: one-shot pruning, which involves…

physics.comp-ph2019

Predicting charge density distribution of materials using a local-environment-based graph convolutional network

Sheng Gong, Tian Xie, Taishan Zhu +4

Electron charge density distribution of materials is one of the key quantities in computational materials science as theoretically it determines the ground state energy and practic…

cs.CV2025

Tera-MIND: Tera-scale mouse brain simulation via spatial mRNA-guided diffusion

Jiqing Wu, Ingrid Berg, Yawei Li +2

Holistic 3D modeling of molecularly defined brain structures is crucial for understanding complex brain functions. Using emerging tissue profiling technologies, researchers charted…

cs.CV2025

The Tenth NTIRE 2025 Efficient Super-Resolution Challenge Report

Bin Ren, Hang Guo, Lei Sun +143

This paper presents a comprehensive review of the NTIRE 2025 Challenge on Single-Image Efficient Super-Resolution (ESR). The challenge aimed to advance the development of deep mode…

cs.CV2023

TinyTracker: Ultra-Fast and Ultra-Low-Power Edge Vision In-Sensor for Gaze Estimation

Pietro Bonazzi, Thomas Ruegg, Sizhen Bian +2

Intelligent edge vision tasks encounter the critical challenge of ensuring power and latency efficiency due to the typically heavy computational load they impose on edge platforms.…

cs.CV2026

Beyond GSD-as-Token: Continuous Scale Conditioning for Remote Sensing VLMs

Song Zhang, Yanlong Chen, Yilin Li +4

Remote sensing vision-language models (RS-VLMs) face a fundamental mismatch with natural-image counterparts: the same geographic object exhibits radically different visual evidence…

cs.RO2026

Robotic Grasping and Placement Controlled by EEG-Based Hybrid Visual and Motor Imagery

Yichang Liu, Tianyu Wang, Ziyi Ye +4

We present a framework that integrates EEG-based visual and motor imagery (VI/MI) with robotic control to enable real-time, intention-driven grasping and placement. Motivated by th…

cs.CL2025

Towards Extreme Pruning of LLMs with Plug-and-Play Mixed Sparsity

Chi Xu, Gefei Zhang, Yantong Zhu +4

N:M structured pruning is essential for large language models (LLMs) because it can remove less important network weights and reduce the memory and computation requirements. Existi…

eess.IV2021

Explaining COVID-19 and Thoracic Pathology Model Predictions by Identifying Informative Input Features

Ashkan Khakzar, Yang Zhang, Wejdene Mansour +5

Neural networks have demonstrated remarkable performance in classification and regression tasks on chest X-rays. In order to establish trust in the clinical routine, the networks'…

cs.LG2024

AttributionLab: Faithfulness of Feature Attribution Under Controllable Environments

Yang Zhang, Yawei Li, Hannah Brown +5

Feature attribution explains neural network outputs by identifying relevant input features. The attribution has to be faithful, meaning that the attributed features must mirror the…

cs.CV2020

Cluster, Split, Fuse, and Update: Meta-Learning for Open Compound Domain Adaptive Semantic Segmentation

Rui Gong, Yuhua Chen, Danda Pani Paudel +5

Open compound domain adaptation (OCDA) is a domain adaptation setting, where target domain is modeled as a compound of multiple unknown homogeneous domains, which brings the advant…

cs.CV2024

Empowering Image Recovery_ A Multi-Attention Approach

Juan Wen, Yawei Li, Chao Zhang +3

We propose Diverse Restormer (DART), a novel image restoration method that effectively integrates information from various sources (long sequences, local and global regions, featur…

cs.CV2026

ATD: Improved Transformer with Adaptive Token Dictionary for Image Restoration

Leheng Zhang, Wei Long, Yawei Li +3

Recently, Transformers have gained significant popularity in image restoration tasks such as image super-resolution and denoising, owing to their superior performance. However, bal…

cs.CV2025

Fractal-IR: A Unified Framework for Efficient and Scalable Image Restoration

Yawei Li, Bin Ren, Jingyun Liang +5

While vision transformers achieve significant breakthroughs in various image restoration (IR) tasks, it is still challenging to efficiently scale them across multiple types of degr…

cs.CV2024

Restore Anything Model via Efficient Degradation Adaptation

Bin Ren, Eduard Zamfir, Zongwei Wu +6

With the proliferation of mobile devices, the need for an efficient model to restore any degraded image has become increasingly significant and impactful. Traditional approaches ty…

cs.LG2022

Machine Learning Applications in Lung Cancer Diagnosis, Treatment and Prognosis

Yawei Li, Xin Wu, Ping Yang +2

The recent development of imaging and sequencing technologies enables systematic advances in the clinical study of lung cancer. Meanwhile, the human mind is limited in effectively…

cs.CV2019

3D Appearance Super-Resolution with Deep Learning

Yawei Li, Vagia Tsiminaki, Radu Timofte +2

We tackle the problem of retrieving high-resolution (HR) texture maps of objects that are captured from multiple view points. In the multi-view case, model-based super-resolution (…

cs.CV2026

Q-MambaIR: Accurate Quantized Mamba for Efficient Image Restoration

Yujie Chen, Haotong Qin, Zhang Zhang +3

State-Space Models (SSMs) have attracted considerable attention in Image Restoration (IR) due to their ability to scale linearly sequence length while effectively capturing long-di…

eess.IV2022

Analyzing the Effects of Handling Data Imbalance on Learned Features from Medical Images by Looking Into the Models

Ashkan Khakzar, Yawei Li, Yang Zhang +5

One challenging property lurking in medical datasets is the imbalanced data distribution, where the frequency of the samples between the different classes is not balanced. Training…

cs.CV2025

IntLoRA: Integral Low-rank Adaptation of Quantized Diffusion Models

Hang Guo, Yawei Li, Tao Dai +2

Fine-tuning pre-trained diffusion models under limited budgets has gained great success. In particular, the recent advances that directly fine-tune the quantized weights using Low-…

eess.IV2024

Q-Segment: Segmenting Images In-Sensor for Vessel-Based Medical Diagnosis

Pietro Bonazzi, Yawei Li, Sizhen Bian +1

This paper addresses the growing interest in deploying deep learning models directly in-sensor. We present "Q-Segment", a quantized real-time segmentation algorithm, and conduct a…

cs.CV2023

Video Super-Resolution Transformer

Jiezhang Cao, Yawei Li, Kai Zhang +1

Video super-resolution (VSR), with the aim to restore a high-resolution video from its corresponding low-resolution version, is a spatial-temporal sequence prediction problem. Rece…

cs.CV2026

The Third Challenge on Image Denoising at NTIRE 2026: Methods and Results

Lei Sun, Hang Guo, Bin Ren +6

This paper reports on the NTIRE 2026 Challenge on Image Denoising, specifically focusing on the high-noise regime (). The competition investigates advanced neural architect…

cs.CV2026

Any Image Restoration via Efficient Spatial-Frequency Degradation Adaptation

Bin Ren, Eduard Zamfir, Zongwei Wu +7

Restoring multiple degradations efficiently via just one model has become increasingly significant and impactful, especially with the proliferation of mobile devices. Traditional s…

cs.LG2025

Calibrating LLMs with Information-Theoretic Evidential Deep Learning

Yawei Li, David Rügamer, Bernd Bischl +1

Fine-tuned large language models (LLMs) often exhibit overconfidence, particularly when trained on small datasets, resulting in poor calibration and inaccurate uncertainty estimate…

eess.IV2021

Plug-and-Play Image Restoration with Deep Denoiser Prior

Kai Zhang, Yawei Li, Wangmeng Zuo +3

Recent works on plug-and-play image restoration have shown that a denoiser can implicitly serve as the image prior for model-based methods to solve many inverse problems. Such a pr…

cs.LG2026

S-CEReBrO: Breaking the Memory Barrier in Continuous EEG Monitoring

Glenn Anta Bucagu, Thorir Mar Ingolfsson, Yawei Li +1

The paper introduces S-CEReBrO, a streaming Transformer architecture that uses a windowed alternating attention mechanism to keep memory usage constant during continuous EEG monito…

#eeg analysis#continuous monitoring#transformer models#attention mechanisms
cs.CL2025

Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs

Hang Guo, Yawei Li, Luca Benini

Recent advances in Large Language Model (LLM) compression, such as quantization and pruning, have achieved notable success. However, as these techniques gradually approach their re…

cs.CV2025

HiM2SAM: Enhancing SAM2 with Hierarchical Motion Estimation and Memory Optimization towards Long-term Tracking

Ruixiang Chen, Guolei Sun, Yawei Li +2

This paper presents enhancements to the SAM2 framework for video object tracking task, addressing challenges such as occlusions, background clutter, and target reappearance. We int…

cs.CV2025

Procedural Generation of 3D Maize Plant Architecture from LIDAR Data

Mozhgan Hadadi, Mehdi Saraeian, Jackson Godbersen +8

This study introduces a robust framework for generating procedural 3D models of maize (Zea mays) plants from LiDAR point cloud data, offering a scalable alternative to traditional…

eess.SP2025

A Compute&Memory Efficient Model-Driven Neural 5G Receiver for Edge AI-assisted RAN

Mahdi Abdollahpour, Marco Bertuletti, Yichao Zhang +3

Artificial intelligence approaches for base-band processing for radio receivers have demonstrated significant performance gains. Most of the proposed methods are characterized by h…

cs.LG2026

Steering Large Reasoning Models towards Concise Reasoning via Flow Matching

Yawei Li, Benjamin Bergner, Yinghan Zhao +3

Large Reasoning Models (LRMs) excel at complex reasoning tasks, but their efficiency is often hampered by overly verbose outputs. Prior steering methods attempt to address this iss…

cs.LG2025

SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models

Wei Huang, Haotong Qin, Yangdong Liu +7

Post-training quantization (PTQ) is an effective technique for compressing large language models (LLMs). However, while uniform-precision quantization is computationally efficient,…

cs.CV2024

Key-Graph Transformer for Image Restoration

Bin Ren, Yawei Li, Jingyun Liang +5

While it is crucial to capture global information for effective image restoration (IR), integrating such cues into transformer-based methods becomes computationally expensive, espe…

eess.SP2025

Finetuning and Quantization of EEG-Based Foundational BioSignal Models on ECG and PPG Data for Blood Pressure Estimation

Bálint Tóth, Dominik Senti, Thorir Mar Ingolfsson +4

Blood pressure (BP) is a key indicator of cardiovascular health. As hypertension remains a global cause of morbidity and mortality, accurate, continuous, and non-invasive BP monito…

cs.CV2026

Efficient Degradation-agnostic Image Restoration via Channel-Wise Functional Decomposition and Manifold Regularization

Bin Ren, Yawei Li, Xu Zheng +6

Degradation-agnostic image restoration aims to handle diverse corruptions with one unified model, but faces fundamental challenges in balancing efficiency and performance across di…

cs.CV2023

Practical Blind Image Denoising via Swin-Conv-UNet and Data Synthesis

Kai Zhang, Yawei Li, Jingyun Liang +6

While recent years have witnessed a dramatic upsurge of exploiting deep neural networks toward solving image denoising, existing methods mostly rely on simple noise assumptions, su…

cs.CV2025

LocalViT: Analyzing Locality in Vision Transformers

Yawei Li, Kai Zhang, Jiezhang Cao +4

The aim of this paper is to study the influence of locality mechanisms in vision transformers. Transformers originated from machine translation and are particularly good at modelli…

cs.CV2019

Learning Filter Basis for Convolutional Neural Network Compression

Yawei Li, Shuhang Gu, Luc Van Gool +1

Convolutional neural networks (CNNs) based solutions have achieved state-of-the-art performances for many computer vision tasks, including classification and super-resolution of im…

cs.CV2025

Ultra-Efficient On-Device Object Detection on AI-Integrated Smart Glasses with TinyissimoYOLO

Julian Moosmann, Pietro Bonazzi, Yawei Li +4

Smart glasses are rapidly gaining advanced functions thanks to cutting-edge computing technologies, especially accelerated hardware architectures, and tiny Artificial Intelligence…

cs.CV2016

Word Recognition with Deep Conditional Random Fields

Gang Chen, Yawei Li, Sargur N. Srihari

Recognition of handwritten words continues to be an important problem in document analysis and recognition. Existing approaches extract hand-engineered features from word images--w…

cs.CV2018

PIRM Challenge on Perceptual Image Enhancement on Smartphones: Report

Andrey Ignatov, Radu Timofte, Thang Van Vu +45

This paper reviews the first challenge on efficient perceptual image enhancement with the focus on deploying deep learning models on smartphones. The challenge consisted of two tra…

cs.CV2024

Hierarchical Information Flow for Generalized Efficient Image Restoration

Yawei Li, Bin Ren, Jingyun Liang +5

While vision transformers show promise in numerous image restoration (IR) tasks, the challenge remains in efficiently generalizing and scaling up a model for multiple IR tasks. To…